Why is a pattern's hit rate meaningless on its own?
A hit rate is meaningless alone because it has no reference point, and the reference point — what happens after an average bar — is usually far closer to the pattern's number than anyone expects.
Suppose a source reports that a bullish engulfing bar is followed by a higher close 56% of the time. That sounds like an edge until you compute the same statistic across all bars in the same period and find it is 55%. The pattern contributed one percentage point, which will not survive the spread.
The reason base rates are rarely quoted is that computing them requires the full dataset rather than the pattern instances, and most informal analysis starts by filtering to the pattern. Once you have filtered, the comparison is no longer available.
So the very first step in any candle study is to compute the outcome for every bar. Everything after that is a comparison, and comparisons are what carry information.
How do you define the outcome you are measuring?
Define the outcome as a symmetric, cost-aware question with a fixed horizon: did price travel X in your favour before travelling X against, within N bars?
That framing is better than 'was the next close higher' for two reasons. It reflects how a trade actually resolves — stops and targets, not closing prices — and it is symmetric, so it cannot be gamed by a horizon that happens to sit in a drifting market.
Express X in ATR rather than pips so the measurement is comparable across instruments and volatility regimes. A twenty-pip target means something different on every pair, and pooling unnormalised results across pairs produces a number that describes the pair mix.
Then include costs in the outcome. Subtract a realistic spread from the favourable threshold and add it to the adverse one. Many candle effects that look real gross of costs disappear entirely net of them, and finding that out is the point of the exercise rather than a disappointment.
How many variants should you count?
Count every definitional choice and every threshold as a separate trial, because candle research typically involves far more of them than the write-up admits.
A modest study of one pattern might involve three body tolerances, two prior-trend requirements, three horizons and two target sizes. That is thirty-six combinations, and the best of thirty-six will look impressive even if the pattern carries no information at all. Reporting only the best combination is the most common way candle statistics become misleading.
The straightforward defence is to record the count and let it set the bar. A result that would be interesting as a first attempt needs to be considerably stronger to be interesting as the best of thirty-six, and there are standard ways to adjust for that once you know the number.
The harder defence is holding a slice of the history back and grading the chosen combination there exactly once. It costs data and it is the only check that cannot be argued with.
What context splits are worth making?
Split by whether the pattern occurred at a structural level, by the prevailing higher-timeframe direction, and by session — because pooling contexts averages an effect that may exist in one with its absence in the others.
Level context tends to matter most in practice. A candle pattern in the middle of a range and the same pattern at a swing high are different events, and the level supplies the selectivity that the candle shape lacks. If you only run one split, run this one.
Direction context matters because reversal patterns and continuation patterns are claims about direction. A bullish engulfing bar in a downtrend and one in an uptrend are testing different propositions, and pooling them tests neither.
Session context matters on instruments with strong intraday liquidity cycles. A pattern forming in a thin session may reflect little participation, and thin-session bars often behave differently for reasons that have nothing to do with the pattern.
Every split costs sample size, which is why they should be chosen deliberately and in advance rather than explored until something interesting appears. Exploring splits is a search over contexts, and the best of ten contexts needs the same scepticism as the best of ten parameters.