Platform · Prove

Proof Engine

A backtest tells you what a strategy would have made. This tells you whether its decisions were actually any good — against history it was never shown.

What you get
Decisions, not an equity curve

Thousands of individual calls, each graded on what happened next.

Held-out by time

Judged on a recent stretch it never trained on. In-sample results are a rehearsal.

It can't see the future

The engine is built so a strategy cannot peek — and every run re-proves it.

A verdict, not a vibe

Deployable, promising, or reject. Including when the answer is reject.

The intrabar order, measured

When one bar holds both your stop and your target, we drop to minute data and settle which came first. Where the minutes are missing it books the loss — never the win.

A finish line for “not yet”

Promising strategies carry the number of forward trades their edge still needs, counted down. A strategy with no positive edge gets no countdown — there is nothing there to prove.

Whatever you build here is tested against real market conditions and graded out-of-sample, so you always know whether it's working — on your own strategy, not a marketing figure. How the proof engine grades.

What we don't publish

We don't publish the grading rule, the split mechanics, or how the no-look-ahead proof is constructed.

We publish what a module does and what it's for. We don't publish how it works. The method is the product — and anyone holding the method holds the product.

Evaluating this for a desk? Talk to sales and we'll go as deep as an NDA allows.

Worth knowing

A verdict is a statement about history, not a forecast. It is the strongest thing we can honestly say and it is still not a promise about next month.

How it works
  1. 01
    Split by time, never by sampling

    History is cut chronologically — fit, select, sealed. The sealed slice is touched exactly once, by the single strategy chosen on the earlier data. Shuffling trades before splitting leaks the future into the past and is how most backtests flatter.

  2. 02
    Ambiguous bars settle against you

    When one bar contains both your stop and your target, finer data decides the order where it exists; where it doesn't, the stop is assumed. We would rather understate an edge than sell a fantasy.

  3. 03
    The search itself is priced in

    Trying many variants and keeping the best is a maximum-of-N statistic — search wide enough over noise and something always looks brilliant. Every sweep reports what the luckiest of N no-edge candidates would have shown.

  4. 04
    One path is not the outcome

    Your equity curve is one ordering of one sample. Resampling reports the drawdown across other orderings of the same trades and the expectancy interval across other draws — because position size gets chosen from those numbers.

Worked example
What the multiple-testing correction actually catches

4,000 sweeps of pure-noise variants, scored two ways. Both tests sound equally rigorous.

Naive test — beat the expected best-of-N
worse than a coin flip at rejecting noise44–53% pass
Tail test — the one we ship
0.9–3.0% pass
Difference
the entire correction

The expected maximum is the centre of the luck distribution, not a threshold. A correction you get wrong is more dangerous than none, because it carries authority.

What it refuses to do

Every one of these is a thing we could ship and choose not to. They are here because the limits are the part of a research tool you actually have to trust.

It won't call a thin sample proven

A winner confirmed on a handful of trades is reported as the beginning of evidence, with the trade count attached.

It won't hide how hard it searched

Every sweep result carries the number of variants tried.

It won't grade a strategy on data it was fitted to

In-sample numbers are labelled a rehearsal and carry no promise.

Where proof engine sits
Describe → Build → Prove → Train → Paper → You arm liveevery improvement re-proves before it counts1DescribePlain English, a pastedscript, or a chart photo2BuildA runnable strategy —one shape everywhere3ProveBacktested and gradedout-of-sample4TrainThe Conscious works it,condition by condition5PaperA real forward record,no money at risk6You arm liveOff by default. Only youcan turn it on

One pipeline, one strategy shape end to end. What you backtest is byte-for-byte what papers and what exports — there is no re-implementation step where drift can hide.

Under the claim
Out-of-sample by timeTen years of history, split by time — never shuffledThe strategy sees thisthe wallThe verdict comesonly from this

A strategy is judged on a recent stretch of history it was never shown. An in-sample curve — the number most backtesters lead with — is a rehearsal, and the platform labels it as one. The engine also re-proves on every run that no rule could see past the wall.

The bar that holds both your stop and your targetOne H1 bar. Your stop AND your target inside it. Which came first?targetstopthe hourly bar can't saySo the engine walks the minutes inside itfirst touch: measuredAnd where the minutes are missing?The loss is booked.Never the win. No exceptions, no flattery.

Most backtesters silently assume the answer that makes the result look better. This platform measures it on minute data — and when the evidence isn't there, it charges you the stop. Your paper trades and your backtests are settled by the same rule.

Research we publish — including the failures
REJECTED
“Fade every liquidity sweep” — the classic smart-money entry, tested raw

≈ 0R over 87,000 sweeps across a decade. Unselected, the famous setup pays nothing.

REJECTED
Volume as an entry gate

Flat across 530,000 samples. We do not gate or size by volume anywhere — and we say so.

REJECTED
Buying dips because the currency is “strong”

No edge over 1,401 graded trades on a decade of data.

VALIDATED
Break-and-retest at a flipped level

A real out-of-sample edge, consistent across every yearly window we held out — fragile to tight stops, which the platform tells you rather than hiding.

VALIDATED
Trailing at the flip wall instead of holding to target

Materially better than holding, measured across 286,000 fade situations.

These are our own studies, run on our own archive, and the negative results ship inside the product next to the positive ones. A platform that only ever finds edges is selling you something. Most ideas don't work; the value is knowing which — before money does the experiment for you.

Questions people actually ask
What does out-of-sample actually mean here?

The engine cuts your history by time into fit, select and sealed slices. Variants are measured on the first, chosen on the second, and the winner is tested exactly once on the third — data nothing was fitted or chosen on.

Why do my results look worse here than on other platforms?

Three reasons, all deliberate: ambiguous bars are settled against you, realistic spread and commission are charged, and the search you ran is priced into the verdict. The number is lower because it is closer to what a broker would have given you.

What is the deflated Sharpe correction?

A haircut for how many strategies you tried before finding this one. Test enough variants on noise and one looks excellent; the correction states what that best-of-N would have scored with no edge at all, so you can see whether yours clears it.

Bring a strategy you already trade.

Test it yourself — or bring your desk's questions to us.

No card required · sales is for desks and teams