← All articles Research / Robustness Testing

The Score That Predicts Which Strategies Will Work Live (103-Strategy Study)

You build a strategy, optimize it, and the backtest looks perfect. Then you take it live and it goes straight down. We wanted to know if a single score, calculated on in-sample data only, could spot the difference in advance. So we built 103 breakout strategies across 8 markets, scored every one with the BreakoutOS Robustness Suite, and checked how they did on data they'd never seen. Of the strategies scoring 65 or higher, 68% stayed positive out-of-sample. Above 80, it was 76.2%.

Why Great Backtests Fail in Live Trading

Most traders have lived through this. The strategy looks beautiful on historical data. The equity curve climbs, the drawdowns are shallow, the profit factor is high. Then it goes live and starts losing money.

The out-of-sample performance doesn't match the in-sample performance. What you were really trading was an overfit: a set of rules tuned so tightly to past data that it had no edge left for the future.

The hard part is that an overfit and a real edge look almost identical on a backtest. Both show a nice equity curve. The question every systematic trader needs answered is how to tell them apart before risking money.

What the BreakoutOS Robustness Suite Measures

The Robustness Suite is a recent addition to BreakoutOS, the platform built for one job: creating breakout strategies. It sits as a layer on top of whatever platform you already trade on, whether that's TradeStation, NinjaTrader, MetaTrader, or TradingView.

Instead of relying on one test, the suite attacks each strategy from several angles. It runs multiple robustness and stress-testing procedures, scores each one, then combines them with a proprietary algorithm into one final score from 0 to 100. That score comes with a verdict: acceptable or not acceptable.

ScoreRatingWhat it means
0-44WeakNo meaningful edge
45-64FragileBelow the recommended threshold
65-84PassedRecommended minimum met
85-100StrongThe highest-scoring strategies

For a closer look at the individual robustness checks, see How BreakoutOS Validates Strategy Robustness.

The Study: 103 Strategies, 8 Markets, Scored Blind

A score is only useful if it predicts something. So we tested whether it does.

We built 103 completely different breakout strategies across 8 markets:

  • US index futures. E-mini S&P 500, E-mini NASDAQ 100, E-mini S&P MidCap 400, E-mini Dow Jones, and E-mini Russell 2000.
  • International index. Nikkei 225.
  • Commodities. Gold.
  • Forex. EUR/GBP.

Each strategy was scored using in-sample data only. The score never saw the out-of-sample period. Then we ran every strategy on the unseen data and compared how different score thresholds lined up with positive, tradable out-of-sample results.

Why this design matters

If the score had access to the test data, a high pass rate would prove nothing. Scoring on in-sample data alone is what makes the out-of-sample results a fair test of whether the score can predict live performance.

The 65 Threshold: 68% vs 48%

BreakoutOS sets the minimum recommended score at 65. Anything 65 or higher shows up green: the strategy has passed.

Here's what that threshold did on unseen data:

In-sample scorePositive out-of-sample
Below 6548%
65 and above68%
Above 8076.2%

Below 65, fewer than half the strategies produced anything meaningful out-of-sample, and the results overall were much worse. At 65 and above, almost 70% delivered positive, tradable results on data the score had never touched.

The rule to remember

The minimum threshold alone lifted the out-of-sample success rate from 48% to 68%. Filtering out strategies below 65 removes a large share of the overfits before they ever reach a live account.

Higher Scores, Higher Pass Rates (76.2% Above 80)

The more interesting finding came when we pushed the threshold higher. The higher the score, the more strategies held up out-of-sample.

For strategies scoring above 80, 76.2% passed on unseen data. In an industry where anything above 60% is already a real improvement, getting close to 80% is about as good as it gets.

One honest caveat: not every score band had a large enough sample to stand on its own. Some individual score intervals didn't have a viable sample size, so read those bands with care. The overall direction is clear, though. Higher scores meant a better chance of surviving unseen data.

See BreakoutOS in Action

Watch how strategies are built, scored, and validated before they ever go live.

Watch Demo Videos  →

Not Every Market Behaved the Same

The results weren't uniform across markets, and the differences are worth understanding.

  • E-mini NASDAQ and E-mini S&P. The strongest markets in the study, with verification rates of 90% or more on true out-of-sample data.
  • EUR/GBP. One of the two weakest. It's a mean-reverting market, which works against breakout strategies.
  • E-mini Russell 2000 (RTY). The other weak spot. Every strategy in the study was long-only, and RTY's behavior didn't suit long breakouts.

The EUR/GBP result is a useful cross-check. In a separate analysis, the Breakout Radar module, which scores how suitable a market is for breakout trading, labeled EUR/GBP as not appropriate for breakout trading at all. The out-of-sample failures here line up with that verdict.

See how markets are scored for breakout suitability in How to Find the Perfect Market for Breakout Trading.

What the Study Didn't Test (and Why It Matters)

These numbers come from the automated score alone. Two steps from a full development process were left out on purpose:

  1. Cross-market validation. Checking whether a strategy's logic also holds up on related markets is a strong extra filter. Adding it would likely push the pass rates higher. See Cross-Market Validation: The Final Test Before Going Live.
  2. Manual review. In our hedge fund, every strategy gets a discretionary review before it trades. That step catches strategies you wouldn't trade even from the in-sample results alone.

So treat the 68% and 76.2% figures as a floor. They show what the score does on its own, before you add the rest of the process.

How to use a robustness score in your own process

  • Treat 65 as a hard gate. Below it, the odds of a strategy holding up were worse than a coin flip in this study.
  • Prefer higher scores when you have a choice. Strategies above 80 held up on unseen data more than three times out of four.
  • Check the market first. A great score on a market that doesn't suit breakouts, like EUR/GBP, is still fighting the market.
  • Keep the human in the loop. Use the score to filter, then review what's left before going live.

Frequently Asked Questions

You can never know for certain, but you can stack the odds. Score the strategy on in-sample data with multiple robustness and stress tests, then confirm it on out-of-sample data it has never seen. In a BreakoutOS study of 103 breakout strategies, 68% of strategies that scored 65 or higher on in-sample data stayed positive out-of-sample, compared with 48% of those below 65. Cross-market validation and a manual review add further protection.
In BreakoutOS, 65 is the recommended minimum, and 85 to 100 is the strong range. In testing on 103 strategies across 8 markets, strategies scoring above 80 held up on unseen data 76.2% of the time. Scores from 45 to 64 are rated fragile, and anything below 45 shows no meaningful edge.
The most common reason is overfitting. When a strategy is optimized too tightly to historical data, it learns the noise of that period instead of a repeatable edge. The backtest looks excellent, but the rules have nothing left to exploit on new data. Trading a market that does not suit the strategy type, such as running breakout strategies on a mean-reverting market, is another frequent cause.
Out-of-sample testing means evaluating a strategy on data that was not used to build, optimize, or score it. The in-sample period is where the strategy is developed; the out-of-sample period acts as a stand-in for live trading. If performance collapses out-of-sample, the strategy was most likely overfit.
The data says it helps. In the BreakoutOS study, applying a minimum robustness score of 65 raised the share of strategies with positive out-of-sample results from 48% to 68%, and requiring a score above 80 raised it to 76.2%. Those figures come from the automated score alone, without cross-market validation or manual review, so a full process should do better still.
Tomas Nesnidal

About the Author

Tomas Nesnidal, known to the systematic trading community as Mr. Breakouts, is a breakout trading specialist, hedge fund co-founder, and creator of BreakoutOS. He has managed institutional portfolios using breakout strategies for over 15 years, trading from 65+ countries. He is the author of The Breakout Trading Revolution and co-founder of Breakout Trading Academy.