How a Backtest Earns Trust
A backtest is a claim about the future dressed up as a fact about the past. It is the easiest thing in this field to produce and the easiest to fool yourself with. So the interesting question isn’t whether a strategy has a good backtest — almost everything that gets published does. The question is what happens when the backtest turns out to be wrong, and who finds out first. This is an account of a number that went up on this site, came down again, and what was learned in between.
The number that looked solid
Until recently the performance page carried a win rate and a Sharpe ratio for the short-dated index strategy, both computed over roughly eight years of simulated history. Several hundred trading days. By the usual standards that is a large sample, and the figures were flattering. They were also, in a way that took some work to see, answering a question nobody had asked.
A fixed distance is not a fixed bet
The strategy places its positions a set distance away from wherever the market is trading. That distance is denominated in index points, and index points are not a stable unit. When the index sat near three thousand, that distance represented a large, comfortable buffer. At current levels the same number of points is a fraction of the move it once implied. Nothing in the rules changed. The market moved underneath them, and the rules quietly came to mean something else.
Which means an eight-year average is not the average of one strategy. It is a blend of a very cautious strategy that rarely lost, a moderately aggressive one, and a distinctly aggressive one, weighted by nothing more principled than how many trading days happened to fall in each era. Split the same history by how far the positions really sat from the market and the loss rates separate cleanly, in order, with no overlap. The single headline figure was an average across three different animals.
There was a tell, too, visible on the chart itself. The distribution of daily outcomes had been drawn with its horizontal scale set by the single worst day in the record — a day from the earliest, least representative period. That one observation stretched the axis so far that the overwhelming majority of days were crushed into a handful of bars. The picture was technically accurate and practically misleading, which is the most dangerous combination a chart can have.
The fix that broke the same way
The obvious correction was to stop simulating and use live results instead. Real fills, real costs, no modelling assumptions to argue about. That went up, and it was wrong for precisely the same reason the first version had been.
The live record also spans more than one strategy. The entry schedule was compressed partway through. A hedging rule — the thing that governs how bad the bad days are allowed to get — was added later still. Averaging across all of it produces a statistic describing something that was never traded in that form for a single day. The error had simply moved down a level, from mixing market regimes to mixing software versions, and it was harder to see there because the word “live” carries an authority that “simulated” does not.
Restricting the sample to only the configuration actually running today fixed the honesty problem and created a different one. What remains is a few dozen trading days. That is enough to know the strategy is working. It is nowhere near enough to characterise how it behaves on its worst days, which is the entire thing a risk statistic is supposed to tell you.
Records and estimates are different objects
The resolution came from separating two things that look similar on a page and are not remotely alike.
A cumulative performance curve is a record. It reports what happened, in sequence. The fact that the strategy was revised along the way is part of that history, not a contaminant in it, and showing the whole thing is straightforwardly honest.
A win rate, a Sharpe ratio, a distribution of outcomes — these are estimates. Each one is a claim that the days in the sample are draws from a single stable process, and that watching enough of them tells you what the next one will look like. Mix in days generated by a different process and the estimate doesn’t become noisy. It becomes a description of nothing at all.
So the curve stayed on the performance page and the summary statistics came down. Not because the results were disappointing — they weren’t — but because there is no sample they could honestly be computed over yet.
When an assumption outvotes the data
One more finding is worth recording, because it is the clearest argument against trusting a simulation that has never been checked against a fill. Reconstructing the hedged results from historical data requires assuming what it would have cost to trade out of a position under pressure. Vary that single cost assumption across a plausible range, holding every piece of market data fixed, and the resulting win rate swings across most of the interval it could possibly occupy. The worst day moves by more than an order of magnitude.
A simulation whose output depends more on one unverifiable input than on eight years of market history is not evidence. It is the assumption, restated at length. The only way to settle that question is to trade it and look at the fills, which is why the figures on this site will lag the strategy rather than lead it.
What has to be true before they go back up
The condition for republishing isn’t a date or a day count. Both would be arbitrary, and a day count in particular is easy to satisfy while still being badly wrong — daily results here are lopsided, with many small gains and occasional large losses, so the sample can look respectably large while the part that actually matters remains barely observed.
Instead the test is a precision one, applied to the current configuration only: resample the live results repeatedly and ask how widely each statistic wanders. While a plausible range for the Sharpe ratio still spans almost everything from unremarkable to implausible, quoting any single value from inside it would be theatre. When that range tightens to something a reader could act on, the numbers go back up. It is checked automatically, and it is not close yet.
The part that costs something
It would be tidier to report that careful analysis revealed the strategy was better than advertised. It didn’t. Every correction here moved the published figures in the unflattering direction, and the most defensible version of the numbers is the least impressive one that has appeared on this site. The honest sample is smaller, the honest win rate is lower, and the honest answer to how it behaves in a crisis is that there isn’t enough evidence yet to say.
That is what a backtest earning trust actually looks like from the inside. Not a clean equity curve, which anyone can produce, but a documented trail of the times the numbers were wrong, how that was discovered, and what was taken down as a result. A track record with no corrections in it isn’t a sign of rigour. It is usually a sign that nobody has looked hard enough yet.