Proof Engine
A backtest tells you what a strategy would have made. This tells you whether its decisions were actually any good — against history it was never shown.
Thousands of individual calls, each graded on what happened next.
Judged on a recent stretch it never trained on. In-sample results are a rehearsal.
The engine is built so a strategy cannot peek — and every run re-proves it.
Deployable, promising, or reject. Including when the answer is reject.
When one bar holds both your stop and your target, we drop to minute data and settle which came first. Where the minutes are missing it books the loss — never the win.
Promising strategies carry the number of forward trades their edge still needs, counted down. A strategy with no positive edge gets no countdown — there is nothing there to prove.
Whatever you build here is tested against real market conditions and graded out-of-sample, so you always know whether it's working — on your own strategy, not a marketing figure. How the proof engine grades.
We don't publish the grading rule, the split mechanics, or how the no-look-ahead proof is constructed.
We publish what a module does and what it's for. We don't publish how it works. The method is the product — and anyone holding the method holds the product.
Evaluating this for a desk? Talk to sales and we'll go as deep as an NDA allows.
A verdict is a statement about history, not a forecast. It is the strongest thing we can honestly say and it is still not a promise about next month.
- 01Split by time, never by sampling
History is cut chronologically — fit, select, sealed. The sealed slice is touched exactly once, by the single strategy chosen on the earlier data. Shuffling trades before splitting leaks the future into the past and is how most backtests flatter.
- 02Ambiguous bars settle against you
When one bar contains both your stop and your target, finer data decides the order where it exists; where it doesn't, the stop is assumed. We would rather understate an edge than sell a fantasy.
- 03The search itself is priced in
Trying many variants and keeping the best is a maximum-of-N statistic — search wide enough over noise and something always looks brilliant. Every sweep reports what the luckiest of N no-edge candidates would have shown.
- 04One path is not the outcome
Your equity curve is one ordering of one sample. Resampling reports the drawdown across other orderings of the same trades and the expectancy interval across other draws — because position size gets chosen from those numbers.
4,000 sweeps of pure-noise variants, scored two ways. Both tests sound equally rigorous.
- Naive test — beat the expected best-of-N
- worse than a coin flip at rejecting noise44–53% pass
- Tail test — the one we ship
- 0.9–3.0% pass
- Difference
- the entire correction
The expected maximum is the centre of the luck distribution, not a threshold. A correction you get wrong is more dangerous than none, because it carries authority.
Every one of these is a thing we could ship and choose not to. They are here because the limits are the part of a research tool you actually have to trust.
A winner confirmed on a handful of trades is reported as the beginning of evidence, with the trade count attached.
Every sweep result carries the number of variants tried.
In-sample numbers are labelled a rehearsal and carry no promise.
One pipeline, one strategy shape end to end. What you backtest is byte-for-byte what papers and what exports — there is no re-implementation step where drift can hide.
A strategy is judged on a recent stretch of history it was never shown. An in-sample curve — the number most backtesters lead with — is a rehearsal, and the platform labels it as one. The engine also re-proves on every run that no rule could see past the wall.
Most backtesters silently assume the answer that makes the result look better. This platform measures it on minute data — and when the evidence isn't there, it charges you the stop. Your paper trades and your backtests are settled by the same rule.
≈ 0R over 87,000 sweeps across a decade. Unselected, the famous setup pays nothing.
Flat across 530,000 samples. We do not gate or size by volume anywhere — and we say so.
No edge over 1,401 graded trades on a decade of data.
A real out-of-sample edge, consistent across every yearly window we held out — fragile to tight stops, which the platform tells you rather than hiding.
Materially better than holding, measured across 286,000 fade situations.
These are our own studies, run on our own archive, and the negative results ship inside the product next to the positive ones. A platform that only ever finds edges is selling you something. Most ideas don't work; the value is knowing which — before money does the experiment for you.
What does out-of-sample actually mean here?
The engine cuts your history by time into fit, select and sealed slices. Variants are measured on the first, chosen on the second, and the winner is tested exactly once on the third — data nothing was fitted or chosen on.
Why do my results look worse here than on other platforms?
Three reasons, all deliberate: ambiguous bars are settled against you, realistic spread and commission are charged, and the search you ran is priced into the verdict. The number is lower because it is closer to what a broker would have given you.
What is the deflated Sharpe correction?
A haircut for how many strategies you tried before finding this one. Test enough variants on noise and one looks excellent; the correction states what that best-of-N would have scored with no edge at all, so you can see whether yours clears it.
Draw your rules on a canvas, or just describe them in plain English. Either way you end up with a strategy the platform can test, trade and grade like any other.
Bring a strategy you already run. An MT5 expert or a Pine script comes in and gets held to the same standard as everything else here.
It reads the chart the way you do — structure, swings, ranges, compression, the shape of the candles — and it reads it as of that bar, never with hindsight.
Most strategies aren't good or bad — they're good somewhere and bad somewhere else. This finds which conditions carry yours, and which quietly bleed it.
Every losing trade is evidence. This reads all of them, finds what they had in common, proposes a change to the rule — and proves the change before it ships.
Every module has an opinion. This turns them into one call, with the reasoning attached — and it will only act on an edge that has actually proved out.
It goes to paper the moment it earns it. Live stays off until you turn it on — and you're the only one who can.
A decade of minute-resolution history across thirty instruments — the thing that decides whether a backtest is evidence or an opinion with a chart attached.
Edges decay. A strategy exported six months ago is quietly rotting on someone's terminal, and nothing tells them. This does.
Bring a strategy you already trade.
Test it yourself — or bring your desk's questions to us.
No card required · sales is for desks and teams