Research · Data study

Does breakout trading actually work? A 23-year, survivorship-free backtest

A Qullamaggie-style breakout detector (leaders near their highs that tighten into a range) tested across 23 survivorship-free years (2004-2026) on 51,126 trades. It flags the breakouts that actually ran — the names like DDD, NVAX, IOVA, and AMBA that ran 5R–50R+ from the breakout line — at their real breakout dates, and shows a statistically significant edge driven entirely by a fat right tail. The honest numbers are smaller than the ones that circulate online, and that is the point.

By Evan Maus~6 min read
51,126
Trades, survivorship-free
23 yrs
Point-in-time, 2004-2026
+$265
Made per $100 risked (mean R)
17 of 17
Years profitable (Kelly, 2010-26)

The short version

  • Across 51,126 trades, every $100 risked made $265 on average (mean R +2.65R, 95% CI [+2.55, +2.75]). Only 36% of trades win. The average winner runs +9.95R; the average loser costs −1.40R.
  • Compounded over 2010-2026 (no leverage, total exposure capital-capped so it never exceeds equity), the equity curve ran from ~65%/yr under flat equal-weight sizing to ~72%/yr sized by optimal fractional-Kelly (each trade weighted by its predicted-R score). Both ends are a backtest ceiling, not a live expectation — small-cap capacity and real fills cap them well below, and the fact that the number swings this much with sizing is exactly why the per-trade R is the honest unit. The Sharpe (1.63) measures return per unit of volatility.
  • The strategy held out-of-sample: the training window (2004-2015) produced mean R +2.52R, and the held-out test window (2016-2026) produced +2.72R. The held-out window held up, if anything stronger, not curve-fitting.
Background

The detector was broken. Here is how it was fixed.

An earlier version of this scanner measured momentum from a stock’s 3-month low with no proximity-to-high gate, so it surfaced laggards. One example sat 28% below its 52-week high when it was flagged. It missed the biggest winners entirely.

The refined detector only considers leaders near their highs that have paused into a tight, well-behaved range, and treats the move out of that range as the entry. It was validated against well-documented momentum winners, including Zoom, NIO, Roku, Tesla, Sea, Enphase, CrowdStrike, and Carvana, which it now flags at their real breakout dates.

The other fix was realism. The prior study used end-of-day prices. The refined study models the fill at the breakout with slippage and accounts for gaps, so a gap through the entry never produces an unrealistic fill. One honest caveat: the headline R assumes the entry is taken right at the line, which rewards precise or automated execution. A trader who waits for confirmation or chases the move captures materially less of it.

Methodology

How the study was constructed

Four choices that make the result harder to manufacture:

  • Survivorship-free universe. The stock universe is rebuilt as it existed on each historical date. Companies that delisted, went bankrupt, or faded into illiquidity stay in the panel.
  • Realistic fills. Entries are modeled at the breakout with slippage rather than an idealized price, and a gap through the intended entry fills at the gap open rather than the pre-gap level.
  • Drawdown tracking. Drawdown is measured on the running equity curve, the figure that matters for real account survivability, not a flattering closed-trades-only number.
  • Risk-normalized returns. The portfolio engine uses fractional-Kelly, score-weighted sizing that is capital-capped (total notional never exceeds equity — no leverage). Results are reported as R (profit/loss in multiples of initial risk) so they are comparable across different account sizes.

The backtest is implemented in the project’s own research code, and every number below is reproduced directly from its output.

Finding 1

The edge is real. It requires patience.

Across 51,126 trades the average result is +2.65R. For every $100 risked, you made $265. The 95% confidence interval is [+2.55, +2.75] and excludes zero.

MetricValue
Trades51,126
Win rate35.7%
Mean R (expectancy)+2.65R
95% CI on mean R[+2.55, +2.75]
Avg winner+9.95R
Avg loser−1.40R
Profit factor3.95
Sharpe ratio (annualized)1.63
Annual return (no leverage: flat → optimal Kelly)65%–72%
Max drawdown (no-leverage Kelly)-19.0%
Per-trade: 2004-2026, entry-at-line, 51,126 trades. CAGR: flat-to-Kelly sizing, no leverage, 2010-2026 (all setups).

The win rate is 36%. The typical trade fails. The strategy makes money because the average winner (+9.95R) runs several times further than the average loser costs (−1.40R), producing a profit factor of 3.95. If a high hit rate is required to stay in the trade psychologically, this structure will grind you down before the edge accumulates.

The +2.65R mean is a fat-tail number, not a smooth per-trade gain. The median trade is a −1R loss: you lose roughly $100 on a $100-risked trade about two of every three times. Roughly 16% of trades exceed +5R and about 10% exceed +10R. Strip those “monster” trades entirely and the mean drops to approximately −0.09R — the entire edge is concentrated in the right tail. That is what the profit factor (3.95) is capturing. Position sizing and not cutting winners early matter far more than win rate.

Compounded over the 2010–2026 simulation window (no leverage, capital-capped), the portfolio ran from 65% CAGR under flat equal-weight sizing to 72% under optimal fractional-Kelly. Both ends are not a live expectation: small-cap capacity and real-world fills cap live returns well below the model. The compounded number swings this much on sizing alone, which is exactly why the per-trade R — sizing-invariant — is the honest unit. The optimal-Kelly curve had a Sharpe of 1.63 and a max drawdown of 19% (no leverage).

2010
+78.0%
2011
+34.0%
2012
+25.0%
2013
+115.0%
2014
+47.0%
2015
+12.0%
2016
+108.0%
2017
+67.0%
2018
+73.0%
2019
+51.0%
2020
+142.0%
2021
+63.0%
2022
+22.0%
2023
+84.0%
2024
+164.0%
2025
+143.0%
2026
+43.0%
Portfolio % return by year under optimal fractional-Kelly sizing (no leverage). 17 of 17 years positive (2010-2026). 2026 is partial.
Finding 2

It held out of sample

The rules were selected on the training window (2004-2015) and run once on the held-out test window (2016-2026). The test window includes 2020, 2022, and 2023, which each tested the strategy in different ways.

WindowMean RNotes
Train (2004-2015)+2.52RRules selected on this window
Test (2016-2026)+2.72RHeld out, rules never saw this data
Walk-forward OOS split: the held-out window held up, no decay.

The test result (+2.72R, or $272 per $100 risked) is, if anything, higher than the training result (+2.52R). The edge did not decay on the held-out years. What matters is the test mean R stays well above zero; curve-fitting would produce a collapse out of sample, not this.

Finding 3

Consistency across two decades

All 17 years in the sim window (2010-2026) were green under optimal Kelly sizing. The leanest were 2015 (+12.0%), 2022 (+22.0%), and 2012 (+25.0%). The best years (2024: +164.0%, 2025: +143.0%) were driven by a handful of monster winners that ran far beyond the initial target — the same fat tail that carries the per-trade mean.

2022 is the tell: a roughly -20% year for the S&P, yet the strategy still returned +22%. That is a property of the near-highs requirement — stocks still holding tight ranges near their highs in a down tape are the handful with genuine relative strength, and a subset break out and run.

Read this part

What this is, and what it isn't

  • Fills are modeled, not guaranteed. Entries assume you get filled at the breakout level, with realistic per-trade slippage and gaps filling at the open, which beats end-of-day pricing but is still optimistic. There are no commissions, spread, or market-impact costs. Real trading shaves these figures, especially in thin names; read the headline R as clean execution, not a promise.
  • CAGR is a Kelly-sizing ceiling, not a forecast. The same no-leverage portfolio compounds at ~65%/yr under flat equal-weight sizing and ~72%/yr under optimal fractional-Kelly (predicted-R weighted, up to 20% per name, capital-capped so total notional never exceeds equity). Both are a model ceiling, not a live expectation: small-cap capacity and real fills cap them well below, and the per-trade edge is unchanged across the range. That the number more than doubles on sizing alone is the point — the per-trade R is the sizing-invariant unit. Max drawdown was 19% on the optimal-Kelly equity curve.
  • 23 years is a solid sample, not a guarantee. The confidence interval on mean R is [+2.55, +2.75], excluding zero. The edge is statistically real over this period; it does not mean it will persist. This is research and education, not financial advice.
  • The detector is specific. The near-highs and tight-range requirements filter out most breakout-looking patterns. A generic momentum screen without these constraints will produce different, probably weaker, results.
So what

What the data informs

Appendix

Methodology and data provenance

Universe: survivorship-free, point-in-time, rebuilt per year. The detector requires leaders near their highs that have tightened into a contained range, with the entry on the move out of that range. Fills are modeled at the breakout with realistic per-trade slippage and account for gaps. Position sizing: fractional-Kelly (predicted-R weighted), capital-capped so total notional never exceeds equity (no leverage), 20% per name, max 10 concurrent positions. R = profit/loss in multiples of initial risk. OOS split: train 2004-2015, test 2016-2026. CI: 95% bootstrap on mean R. 51,126 trades, 2004-2026. Study generated July 2026. This is research and education, not investment advice.

The edge is mechanical, and mechanical means trainable.

The data rewards a repeatable process: pick only the bases that qualify, size by risk, cut losers fast, let winners run. The drill shows a real setup, makes you call it, then shows what happened.

Does Breakout Trading Work? A 23-Year Backtest