Key takeaways · 6
- Every S&P 500 correction of 10% or more between 1996 and 2026 — all 11 of them, including 2002, 2009, 2020, and 2022 — produced a Follow-Through Day within 30 sessions of its final low.
- Across all 49 FTD signals a rule-follower would have acted on, 51% eventually closed back below their rally low. The signal has perfect recall of major bottoms — every 10%+ correction produced one — and roughly coin-flip precision.
- Three months after an FTD the S&P 500 was up a median +4.4% versus +3.3% from any in-correction session — about a one-point edge that 30 years of data cannot cleanly separate from zero (95% CI on the difference: −5.0 to +4.2 points). The popular telling oversells it.
- Failed FTDs fail fast: half of all failures undercut the rally low within 22 sessions, a quarter within 6. The stop-out does its job early or not at all.
- The higher-volume requirement — the rule's signature clause — added no measurable edge in 30 years of data. The 95% confidence interval on its 3-month benefit runs from −2.3 to +4.1 points.
- FTDs during bear markets of 20%+ were traps until the final low: the first signal of those episodes fell a median −6.6% over the next 3 months and failed half the time.
The two numbers that define the signal
Between January 1996 and May 2026 the S&P 500 went through 11 corrections of 10% or more. Every single one of them — the dot-com bust, the 2008 crisis, the COVID crash, the 2022 bear, the 2025 tariff break — produced a Follow-Through Day within 30 sessions of its final low. Eleven bottoms, eleven FTDs, no exceptions. William O'Neil's famous claim that no bull market has ever started without one holds, at least on three decades of index data.
Here is the other number. Over the same 30 years the rule produced 49 signals a trader following it in real time would have acted on. Twenty-five of them — 51% — eventually closed back below the rally low they confirmed. The signal caught every real bottom by also firing at plenty of fake ones.
Both numbers are true at once, and together they say what a Follow-Through Day actually is: a high-recall, low-precision regime signal. It will not miss the turn. It will also cry wolf about as often as it calls one. The interesting questions are the ones underneath — how much edge the signal carries over just buying any correction day, how quickly the failures reveal themselves, and whether the celebrated details of the rule (the volume requirement, the day-4-to-7 window) actually carry information. That is what this study measures.
All 49 Follow-Through Day signals, 1996–2026. Red dots closed back below their rally low within 63 sessions. Note where they cluster: inside prolonged bears, where each new leg down minted a fresh doomed signal.
How we measured it
Everything below comes from one deterministic pass over S&P 500 index daily closes and composite volume — 7,648 sessions from January 2, 1996 to May 22, 2026, held in the same bar store that feeds the TickerStance dashboard. Index-level composite volume is closer to the exchange volume O'Neil originally used than ETF share volume, and the series has no missing sessions against the Dow series we cross-checked it with.
The detection rule is the classic one, stated so it can be replayed: a correction begins when the index closes 5% or more below its running peak. After each new low of that correction, a rally attempt starts on the first higher close; that session is day 1. A Follow-Through Day is the first session on day 4 or later of a surviving attempt that closes up at least 1.25% on volume above the prior session. A new low voids the attempt and the count restarts. Failure means what O'Neil said it means: the index closes back below the rally low the signal confirmed. We track whether that happens within 21 sessions, within 63, and at any point before the index regains its prior peak.
Sample sizes are what they are: 30 correction episodes, 49 acted-on signals, 11 major bottoms. We report medians with bootstrap 95% confidence intervals — resampling whole correction episodes, not individual days, so the intervals respect the fact that returns cluster and overlap — and quote them wherever a claim depends on them. Nothing here was tuned after seeing the results: the rules, thresholds, outcome definitions, and baselines were fixed first, and any analysis we added after the fact is flagged as exploratory where it appears.
| Methodology item | Value |
|---|---|
| Instrument | S&P 500 index, daily close and composite volume |
| Window | January 2, 1996 – May 22, 2026 (7,648 sessions) |
| Correction definition | Close 5%+ below the running peak close |
| Rally attempt | First up-close after a correction low; voided by a new low |
| FTD rule | Day 4+ of an attempt, gain ≥ 1.25%, volume above the prior session |
| Failure | Any close back below the rally low the signal confirmed |
| Signals (acted-on cohort) | 49, one per rally attempt |
| Correction episodes | 30, of which 11 reached −10% or deeper |
| Inference | Bootstrap 95% CIs, 10,000 resamples, fixed seed |
The base case: what follows an FTD
Start with the honest benchmark question. An FTD can only fire during or just after a correction, so the fair comparison is not the market's average day — it is the average day inside a correction, when you could have simply bought without waiting for any signal.
Three months after the median FTD, the S&P 500 was up 4.4% (95% CI: +1.0 to +6.3), with 67% of signals positive. Three months after the median in-correction session, it was up 3.3% with 65% positive. Twelve months out the FTD cohort shows a median +10.8% (CI: +4.0 to +17.5) against a long-run base rate that makes that number respectable but not miraculous.
Now the honest part, held to the same standard this study applies to the volume rule later. The FTD cohort's median beats the baseline at every horizon, but not by a margin the sample can stand behind. Resampling whole correction episodes, the three-month difference is +1.1 points with a 95% confidence interval of −5.0 to +4.2 points; the one-month difference is +0.7 points, CI −1.9 to +2.5. Both straddle zero. Thirty years of index history cannot establish that buying the Follow-Through Day beats buying any day of the same correction. The point estimates lean the FTD's way and its win rates run about two points higher, so the tilt is directionally there — but it is a tilt the data cannot certify, not the decisive edge the folklore implies.
And note the spread underneath those medians: the interquartile range at three months runs from −2.7% to +7.8%, and the mean (+1.8%) sits well below the median because the left tail is long. The average path hides a wide fan of outcomes. This is the study's first uncomfortable result, and it is worth sitting with: the celebrated part of the Follow-Through Day — that it front-runs a rally — is the part 30 years of data support least.
One more number worth keeping: the median FTD arrived 7 sessions after the low it confirmed, with the index already 5.5% off that low. That is the price of confirmation. The signal will never sell you the bottom tick — it is designed to cost you the first few percent in exchange for evidence.
The median path after an FTD (solid, with interquartile band) against the same statistic from every in-correction session (dashed). The solid line leads the dashed one — but by less than the band is wide, which is the whole caveat.
Failure anatomy: when the signal is wrong, it says so quickly
Twenty-two percent of signals closed back below their rally low within a month. Thirty-nine percent did within three months. Fifty-one percent did eventually, before the index regained its prior peak. Those are the precision numbers, and they are why nobody should size a position as if an FTD were a guarantee.
But look at the timing of the failures rather than just their count. Among the 25 signals that failed, the median time from FTD to the undercut was 22 sessions. A quarter failed within 6 sessions. Three quarters had resolved within 55. The distribution has a short fuse: a Follow-Through Day that is going to be wrong usually shows its hand within a month, and the tell is mechanical — a close below a level you knew the day you entered.
That is the practical redemption of a coin-flip precision rate. A signal that is wrong half the time but tells you quickly and cheaply is usable; a signal that is wrong half the time and lets losses run is not. The framework's stop-out rule is not an accessory to the FTD — it is the half of the system that makes the other half tolerable. (Our failure line, the rally low, sits below the FTD day's own intraday low that IBD practitioners use as a stop, so if anything the classic stop exits earlier than the failures counted here.)
When failed FTDs failed. The mass sits inside the first month — the signal resolves fast in both directions.
Eleven bottoms, eleven Follow-Through Days
Here is the recall side of the ledger in full. For every S&P 500 correction that reached 10% or deeper, the table shows the final low and the first FTD after it. The dates will look familiar to anyone who lived through them. The clearest external check is January 4, 2019 — the +3.4% Powell-pivot session that flipped IBD's own market call to a confirmed uptrend; our detector lands on it without being told to. (We validate the detector more rigorously against our own production pipeline further down, where the replay reproduces every live-flagged FTD date exactly.)
Two honesty notes before the table. First, this cohort is labeled in hindsight: "the FTD after the final low" is only knowable once the low is final, so these rows demonstrate that the signal fires at real bottoms — not that you could have known those particular signals were the real ones. Second, a statistical wrinkle: an FTD that fires after the final low can never "fail" by our definition, because a close below that low would simply have made a new final low. Their perfect record is partly definitional. What is not definitional is that the rule fired at all 11 bottoms, quickly — a median 7 sessions after the low — and that the subsequent year was up a median 21.3% across the cohort.
| Correction | Depth | Final low | First FTD after the low | FTD gain |
|---|---|---|---|---|
| 1997 (Asian crisis) | −10.8% | Oct 27, 1997 | Nov 20, 1997 · day 18 | +1.5% |
| 1998 (LTCM) | −19.3% | Aug 31, 1998 | Sep 8, 1998 · day 5 | +5.1% |
| 1999 (rate scare) | −12.1% | Oct 15, 1999 | Oct 28, 1999 · day 9 | +3.5% |
| 2000–02 (dot-com bust) | −49.1% | Oct 9, 2002 | Oct 15, 2002 · day 4 | +4.7% |
| 2007–09 (financial crisis) | −56.8% | Mar 9, 2009 | Mar 18, 2009 · day 7 | +2.1% |
| 2015–16 (China/oil) | −14.2% | Feb 11, 2016 | Mar 1, 2016 · day 12 | +2.4% |
| 2018 (Volmageddon) | −10.2% | Feb 8, 2018 | Feb 14, 2018 · day 4 | +1.3% |
| 2018 (Q4 slide) | −19.8% | Dec 24, 2018 | Jan 4, 2019 · day 7 | +3.4% |
| 2020 (COVID crash) | −33.9% | Mar 23, 2020 | Apr 2, 2020 · day 8 | +2.3% |
| 2022 (rate bear) | −25.4% | Oct 12, 2022 | Oct 21, 2022 · day 7 | +2.4% |
| 2025 (tariff break) | −18.9% | Apr 8, 2025 | Apr 22, 2025 · day 9 | +2.5% |
The bear-market trap
The failures are not spread evenly through time. They concentrate in one specific situation: the middle of a deep bear market.
Split the correction episodes by depth and look at the first FTD each produced. In shallow corrections (5–10%), the first signal went on to a median +6.0% over three months and failed only 15% of the time. In 10–20% corrections: +5.5% and 29%. In the four bear markets that exceeded 20% — 2000–02, 2007–09, 2020, 2022 — the first FTD of the episode returned a median −6.6% over the next three months and failed half the time.
The mechanism is visible on the timeline chart above, and the 2007–09 bear is its cleanest exhibit: that decline produced eight consecutive failed Follow-Through Days — December 2007, January 2008, the post-Bear-Stearns rally, July, September, twice in October, December — before the ninth, on March 18, 2009, held and ran for years. (Our strict close-only day count dates that Bear Stearns FTD to March 20, 2008; the concept explainer that walks the 2008 failure in detail follows IBD's March 18 count off the same rally — the two-day gap is a counting convention, covered in the limitations below.) Each new leg down voided the old signal and eventually spawned a new one. The FTDs that mattered in the deep bears were the last ones, and nothing in the rule itself distinguishes the last from the first.
The practical reading: an FTD during a young, shallow correction deserves more benefit of the doubt than an FTD sixteen months into a grinding bear with three failed signals behind it. The rule fires the same way both times; the base rates underneath are very different. Prior failed attempts in the same decline are information the rule ignores and you should not.
The volume test: the rule's signature clause doesn't earn its keep
Every statement of the Follow-Through Day rule leans on volume: the gain must come on volume above the prior session, because that is the institutional footprint. It is the clause that separates the FTD from "the market went up a lot." So we tested it directly.
Take every session inside a correction that gained 1.5% or more. Of those, 228 cleared the volume bar and 180 did not. If the volume clause carries information, the first group should outperform the second. Measured over the sessions with three months of subsequent data, the higher-volume group had a median return of +4.9% and the lower-volume group +3.7% — a 1.2-point gap in the expected direction, but with a bootstrap 95% confidence interval on the difference, resampling whole correction episodes, running from −2.3 to +4.1 points. Thirty years of data cannot tell that gap from zero.
A softer version survives as a hint, not a finding: in the robustness grid below, requiring volume above both the prior day and the 50-day average trimmed the failure rate a few points versus no volume condition at all. Directionally friendly to O'Neil, statistically unproven. Our read is that the volume clause is close to free — it filters little and costs little — but traders who agonize over whether an FTD "counts" because volume ran 2% below the prior session are agonizing over noise. It is worth saying plainly, because the folklore treats the volume test as the heart of the signal, and on this evidence it is closer to a garnish.
The day-count test: patience beats the classic window
O'Neil taught that the strongest FTDs arrive on days 4 through 7 of the rally attempt, and some practitioners discount signals after day 10 entirely. Our data disagrees with the cutoff.
We re-ran the full study across a grid of rule variants: gain thresholds from 1.0% to 2.0%, three volume conditions, and the day window capped at day 7 versus left open. These rows use the episode-first cohort — the first FTD of each correction, 24 signals — so the failure rates differ from the 49-signal numbers above. The pattern is consistent everywhere: restricting the window to days 4–7 produced fewer signals with higher failure rates — at our primary threshold, a 47% failure rate within three months versus 25% with the window left open. Eight of the eleven major-bottom FTDs arrived on day 7 or later; the 2016 bottom confirmed on day 12, the 1997 bottom on day 18. A day-7 cutoff throws those away and keeps the early signals that fail more.
This cuts against how the rule is usually taught and against the 2008-era finding from Quantifiable Edges that late FTDs underperform, so treat the disagreement with appropriate humility — different windows, different failure definitions, different eras. But on 1996–2026 S&P 500 data, the conclusion is not close: the day-4 minimum does real work (day-1-to-3 confirmations are noise by construction), and the upper bound of the classic window does not.
| Rule variant (gain ≥ 1.25%) | Signals | Median 3-month return | Failed within 63 sessions |
|---|---|---|---|
| Volume above prior day, days 4–7 only | 15 | +6.4% | 46.7% |
| Volume above prior day, day 4+ | 24 | +5.5% | 25.0% |
| Volume above prior day and 50-day average, days 4–7 only | 13 | +5.9% | 53.8% |
| Volume above prior day and 50-day average, day 4+ | 23 | +6.0% | 21.7% |
| No volume condition, days 4–7 only | 22 | +5.8% | 31.8% |
| No volume condition, day 4+ | 27 | +4.9% | 25.9% |
What separates the winners: distribution days after the signal
O'Neil's own tiebreaker for a fresh FTD was the behavior that follows it: an uptrend that immediately starts logging distribution days — declines of 0.2%+ on rising volume — is an uptrend under sale. We counted them across the 25 sessions after each of the 49 signals. Two caveats up front, because they bound how far this can be pushed. This analysis was added after seeing the failure data, so it is exploratory, not pre-registered. And it is a raw count of down-volume days — it omits the IBD aging rule the dashboard's own distribution_days signal applies (a distribution day is cancelled once price rallies 5% above it), so these counts run higher than the dashboard number and are not the same metric.
The split is large. Signals followed by four or fewer distribution days in their first 25 sessions failed 20% of the time. Signals followed by five or more failed 52% of the time. The counts themselves skew high because post-correction tape is volatile — almost every FTD sees at least three — but the five-plus cluster marked more than half of the eventual failures.
Read causally with care, because the window works against a clean reading: failures land at a median 22 sessions, inside the 25-session count, so a failing FTD mechanically generates distribution days as it drops. Much of the association is the failure already in progress, not an early warning of one. What survives that objection is modest and practical — a fresh FTD that stays quiet on the distribution front is a better bet than one that immediately starts logging heavy selling — and it is the same follow-through-to-enter, distribution-to-exit logic the dashboard's Breadth subscore runs each session, though the dashboard uses the aged signal, not this raw count.
Every signal, every horizon, one dot each. The medians (marked) drift right as the horizon extends; the left tail belongs mostly to the signals that failed within 63 sessions.
What this study cannot tell you
Forty-nine events over thirty years is a real sample for a rare signal, and it is still forty-nine events. The confidence intervals quoted above are wide, and several cuts (the depth split in particular) run on single-digit or low-double-digit counts — we quote them as description, not proof. Overlapping forward windows add another layer: signals that cluster in the same bear share the same subsequent market, so the effective sample is smaller than the nominal one and the intervals are, if anything, too narrow.
One reassurance against the small sample: running the identical detection on a second index — the Nasdaq 100 over the same 1996–2026 window — produced 38 signals with the same shape (a +6.7% median three-month return, a 26% three-month failure rate, +30% at a year). The core picture is not an artifact of the S&P 500 alone. But two correlated US large-cap indexes are not independent evidence, and everything here is one market regime — a thirty-year US bull-biased sample — so treat the numbers as a careful description of that history, not a law.
The detection is close-only — no intraday undercuts, no "closes in the upper half of the range" refinements — and the day count starts at the first up-close after a low, which is one of several counting conventions in circulation. Ours dated the 2009 FTD to March 18; IBD's count, which can start the clock on the low day itself, called March 12. Same bottom, same conclusion, different day label. The composite-volume series also changed character across three decades of market-structure evolution, which is one more reason we lean on the day-over-day volume comparison rather than volume levels.
Most importantly: nothing here is a strategy backtest. There are no positions, no costs, no sizing, and the S&P 500 rose over the window, which flatters every long-only conditional average including the baselines. The study measures what the market did after a signal, against honest baselines of what it did anyway. What a trader does with that tilt — and at what size, with what stop — is a separate problem this page does not solve. None of this is investment advice.
How this connects to the dashboard signal
TickerStance's own follow_through_day signal is a deliberately simplified cousin of the rule studied here: SPY only, a 1.5% gain bar, volume above the prior day, a 5% correction precondition instead of a day count, and a rolling 30-session activation window. We replayed that exact rule over the same 30 years as a cross-check — and reproduced, to the day, all seven FTD dates the production pipeline has flagged since it went live.
The simplification has a measurable cost, and the study quantifies it: without the day-4 waiting period, the production rule confirms earlier and fails more often — 44% of its episode-first signals undercut within three months, versus 25% for the classic day-count rule, with essentially the same median return when it holds. That is a defensible trade for a dashboard signal that must be deterministic and explainable, and it is exactly why the site pairs it with a distribution-day count on the same index: the pairing, not either signal alone, is the regime read. The numbers above are the base rates behind that design choice.
Frequently asked questions
What is the success rate of a follow-through day?
In our study of S&P 500 index data from 1996 to 2026, 49% of follow-through day signals held (the index never closed back below the rally low they confirmed) and 51% eventually failed. Within three months of the signal, 61% were still intact and the median forward return was +4.4%.
Has every bull market really started with a follow-through day?
In our 30-year window, yes: all 11 S&P 500 corrections of 10% or deeper produced a follow-through day within 30 sessions of their final low, including the 2002, 2009, 2020, and 2022 bear-market bottoms. The catch is that the signal also fired dozens of other times that did not mark final lows.
How quickly do failed follow-through days fail?
Fast. Among the 25 failed signals in our sample, half undercut their rally low within 22 trading sessions of the FTD and a quarter within 6. Three quarters of failures resolved within 55 sessions. A follow-through day that is going to be wrong usually says so within a month.
Does the higher-volume requirement actually matter?
Less than the folklore says. Comparing 1.5%+ up days in corrections with and without the volume condition, the higher-volume group led by 1.2 points of median 3-month return, but the 95% confidence interval on that difference (−2.3 to +4.1 points, resampling whole correction episodes) includes zero. Thirty years of data cannot confirm the volume clause adds edge.
Is a follow-through day on day 4 to 7 stronger than a later one?
Not in this data — the opposite. Restricting signals to the classic day-4-to-7 window raised the three-month failure rate from 25% to 47% at our primary threshold, and 8 of the 11 major-bottom FTDs arrived on day 7 or later. The day-4 minimum helps; the day-7 ceiling does not.
Should I buy stocks when a follow-through day fires?
This study measures market behavior after the signal; it is not investment advice and not a strategy backtest. What the data supports: the signal reliably marks major bottoms and flags its own failures quickly via the rally low. What it does not support is a forward-return edge over simply buying the same correction — that difference is not statistically distinguishable from zero at 30 years of sample. Reliability also depends heavily on context: fresh shallow corrections good, deep bears with prior failed signals bad.