You build a strategy, optimize it, and the backtest looks perfect. Then you take it live and it goes straight down. We wanted to know if a single score, calculated on in-sample data only, could spot the difference in advance. So we built 103 breakout strategies across 8 markets, scored every one with the BreakoutOS Robustness Suite, and checked how they did on data they'd never seen. Of the strategies scoring 65 or higher, 68% stayed positive out-of-sample. Above 80, it was 76.2%.
Why Great Backtests Fail in Live Trading
Most traders have lived through this. The strategy looks beautiful on historical data. The equity curve climbs, the drawdowns are shallow, the profit factor is high. Then it goes live and starts losing money.
The out-of-sample performance doesn't match the in-sample performance. What you were really trading was an overfit: a set of rules tuned so tightly to past data that it had no edge left for the future.
The hard part is that an overfit and a real edge look almost identical on a backtest. Both show a nice equity curve. The question every systematic trader needs answered is how to tell them apart before risking money.
What the BreakoutOS Robustness Suite Measures
The Robustness Suite is a recent addition to BreakoutOS, the platform built for one job: creating breakout strategies. It sits as a layer on top of whatever platform you already trade on, whether that's TradeStation, NinjaTrader, MetaTrader, or TradingView.
Instead of relying on one test, the suite attacks each strategy from several angles. It runs multiple robustness and stress-testing procedures, scores each one, then combines them with a proprietary algorithm into one final score from 0 to 100. That score comes with a verdict: acceptable or not acceptable.
| Score | Rating | What it means |
|---|---|---|
| 0-44 | Weak | No meaningful edge |
| 45-64 | Fragile | Below the recommended threshold |
| 65-84 | Passed | Recommended minimum met |
| 85-100 | Strong | The highest-scoring strategies |
For a closer look at the individual robustness checks, see How BreakoutOS Validates Strategy Robustness.
The Study: 103 Strategies, 8 Markets, Scored Blind
A score is only useful if it predicts something. So we tested whether it does.
We built 103 completely different breakout strategies across 8 markets:
- US index futures. E-mini S&P 500, E-mini NASDAQ 100, E-mini S&P MidCap 400, E-mini Dow Jones, and E-mini Russell 2000.
- International index. Nikkei 225.
- Commodities. Gold.
- Forex. EUR/GBP.
Each strategy was scored using in-sample data only. The score never saw the out-of-sample period. Then we ran every strategy on the unseen data and compared how different score thresholds lined up with positive, tradable out-of-sample results.
Why this design matters
If the score had access to the test data, a high pass rate would prove nothing. Scoring on in-sample data alone is what makes the out-of-sample results a fair test of whether the score can predict live performance.
The 65 Threshold: 68% vs 48%
BreakoutOS sets the minimum recommended score at 65. Anything 65 or higher shows up green: the strategy has passed.
Here's what that threshold did on unseen data:
| In-sample score | Positive out-of-sample |
|---|---|
| Below 65 | 48% |
| 65 and above | 68% |
| Above 80 | 76.2% |
Below 65, fewer than half the strategies produced anything meaningful out-of-sample, and the results overall were much worse. At 65 and above, almost 70% delivered positive, tradable results on data the score had never touched.
The rule to remember
The minimum threshold alone lifted the out-of-sample success rate from 48% to 68%. Filtering out strategies below 65 removes a large share of the overfits before they ever reach a live account.
Higher Scores, Higher Pass Rates (76.2% Above 80)
The more interesting finding came when we pushed the threshold higher. The higher the score, the more strategies held up out-of-sample.
For strategies scoring above 80, 76.2% passed on unseen data. In an industry where anything above 60% is already a real improvement, getting close to 80% is about as good as it gets.
One honest caveat: not every score band had a large enough sample to stand on its own. Some individual score intervals didn't have a viable sample size, so read those bands with care. The overall direction is clear, though. Higher scores meant a better chance of surviving unseen data.
See BreakoutOS in Action
Watch how strategies are built, scored, and validated before they ever go live.
Watch Demo Videos →Not Every Market Behaved the Same
The results weren't uniform across markets, and the differences are worth understanding.
- E-mini NASDAQ and E-mini S&P. The strongest markets in the study, with verification rates of 90% or more on true out-of-sample data.
- EUR/GBP. One of the two weakest. It's a mean-reverting market, which works against breakout strategies.
- E-mini Russell 2000 (RTY). The other weak spot. Every strategy in the study was long-only, and RTY's behavior didn't suit long breakouts.
The EUR/GBP result is a useful cross-check. In a separate analysis, the Breakout Radar module, which scores how suitable a market is for breakout trading, labeled EUR/GBP as not appropriate for breakout trading at all. The out-of-sample failures here line up with that verdict.
See how markets are scored for breakout suitability in How to Find the Perfect Market for Breakout Trading.
What the Study Didn't Test (and Why It Matters)
These numbers come from the automated score alone. Two steps from a full development process were left out on purpose:
- Cross-market validation. Checking whether a strategy's logic also holds up on related markets is a strong extra filter. Adding it would likely push the pass rates higher. See Cross-Market Validation: The Final Test Before Going Live.
- Manual review. In our hedge fund, every strategy gets a discretionary review before it trades. That step catches strategies you wouldn't trade even from the in-sample results alone.
So treat the 68% and 76.2% figures as a floor. They show what the score does on its own, before you add the rest of the process.
How to use a robustness score in your own process
- Treat 65 as a hard gate. Below it, the odds of a strategy holding up were worse than a coin flip in this study.
- Prefer higher scores when you have a choice. Strategies above 80 held up on unseen data more than three times out of four.
- Check the market first. A great score on a market that doesn't suit breakouts, like EUR/GBP, is still fighting the market.
- Keep the human in the loop. Use the score to filter, then review what's left before going live.

