What earns a verdict in a forex backtest.
Nine standards every strategy is held to — yours, and every template on the platform. They're the reason the proof engine tells you no more often than it tells you yes.
In-sample results are a rehearsal
Any rule can be made to look good on the data used to shape it. So the number that counts here is measured on a recent stretch of history the strategy was never shown. If it only works where it was fitted, it doesn't work.
EMA Crossover made +0.037R in-sample and lost 0.100R out-of-sample. That's the whole problem, in one row.
Decisions, not equity curves
An equity curve is one path — get lucky in the right month and it flatters everything. We grade the individual decisions instead, thousands of them, each against what actually happened next. A curve can lie about a strategy. A few thousand graded calls can't.
A strategy must never see the future
Most backtest bugs are one form of this — a value that wasn't knowable yet quietly leaking into a decision. Our engine is built so it can't happen, and every run re-proves that rather than trusting the design.
Costs are real, so we charge them
Spread, commission and slippage are applied. A strategy that only survives without costs isn't a strategy, it's a spreadsheet.
Enough evidence, or no verdict
A handful of good trades is a story. We won't call something proven until there's enough of it to be statistically meaningful — and plenty of ideas simply stay unproven.
Measured beats estimated — and we label which is which
Some things we can measure directly: what a change did to real trades. Some we can only estimate, like the losses avoided by never trading a rejected idea. Both are useful. Only one is evidence, and we never quietly add an estimate into a measured headline.
When the stop and the target sit in the same bar, we go and look
Every backtester meets this: one bar touches both your stop and your target, and the order decides whether the trade was a win or a loss. Most tools pick a rule and move on. We drop to minute data and measure which came first. Where the minutes exist, the trade is settled by evidence. Where they don't, it books the loss — never the win, never a coin flip.
Ten years of EUR/USD on the hourly chart: 26 bars held both. Every one of them resolved against the actual minute-by-minute path, and the result names the timeframe that settled it.
The number is never harsher than the evidence, either
Conservative is not the same as correct. A backtest that assumes the worst on a bar we could have measured is understating a real edge — a quieter mistake than overstating one, and just as wrong. So the engine reaches for the finest data that actually covers each bar before it falls back to assuming.
We found this in our own engine: a partial minute archive made ten-year backtests resolve FEWER bars than none at all, because the code committed to minutes it didn't have instead of using the five-minute data it did. 2 bars measured became 26.
“Almost proven” is a number, not a feeling
A forward test needs a finish line, and it depends on the size of the edge — a thin one takes far longer to separate from luck than a strong one. We compute how many trades this specific strategy needs, and count toward it. If the edge isn't positive, there's no countdown at all: telling you a losing strategy is merely unproven would be the most expensive thing we could say.
The resolution the intrabar question is settled at.
FX majors and crosses, indices, metals, oil, crypto.
Back to July 2016 — through 2020, through 2022.
Depth is the point, not the number. A decade spans regimes a two-year window cannot — a strategy that only ever saw 2023 has never been told no by a real market. And minute resolution is what makes standard 07 possible at all: without it, every bar that holds both your stop and your target is a guess dressed as a result.
Every bar is checked against history we already hold before it is allowed in, and a feed that disagrees is refused rather than blended. That check has earned its place: it caught one instrument arriving with every bar timestamped an hour late — data that looked perfectly reasonable in aggregate and would have quietly walked the wrong minutes for every trade tested on it.
Positive out-of-sample, statistically meaningful, and it held up. Rare.
Positive out-of-sample but not yet significant. Interesting, not evidence.
No real out-of-sample edge. Most things land here. Most things should.
Held to this bar, most strategies land on reject — and that's the point. A tool that mostly agrees with you isn't testing anything. When your idea comes back deployable, you'll know it earned it.
Promising is not a holding pen. Each one carries the number of forward trades it still needs before the edge could be called real, counted down as they happen — so “not yet” has a finish line instead of being somewhere ideas go to sit.
We don't publish the grading rule itself, how the held-out split is constructed, how significance is computed, or how the no-look-ahead proof is built.
We publish what a module does and what it's for. We don't publish how it works. The method is the product — and anyone holding the method holds the product.
Evaluating this for a desk? Book a demo and we'll go as deep as an NDA allows.
None of the standards above are secret — anyone serious already knows they're the right bar. Knowing the bar and being able to enforce it are different problems, and the second one is the product.
Put a strategy through it.
Yours, graded against real conditions — including when the answer is no.
No card required · demos are for desks and teams