What is the Turtle Soup strategy?
Turtle Soup is a counter-trend rule that enters against a breakout of a recent extreme when that breakout fails to sustain. Price takes out a high or low from the last several sessions, does not follow through, and the rule enters in the opposite direction with a stop beyond the extreme just made.
The name is a reference to fading the breakout entries that trend-following systems of an earlier generation were known for. The premise is that a break of an obvious level attracts breakout orders and stop orders, and that a break which cannot hold leaves those participants offside, which supplies the fuel for the move back.
Compared with most named patterns this one is refreshingly specific. There is a lookback period defining the extreme, a break condition, and a failure condition. Each is a number rather than a judgement, which means the whole thing can be enumerated mechanically without any labelling step.
That specificity is exactly why it is worth using as a template. A pattern that cannot be reduced to this form usually cannot be graded at all.
How is a failed breakout defined precisely?
A failed breakout needs three numbers: the lookback that defines the extreme, how far beyond it price must travel to count as a break, and what evidence counts as failure within a stated time limit.
The lookback is the least controversial and still a parameter. A twenty-session extreme and a four-session extreme select different events, occurring at different frequencies with different characteristics, and the choice cannot be made by testing several and keeping the best without disclosing that.
The break threshold matters more than it looks. Requiring price merely to touch the prior extreme catches a great many trivial events, including ones that are indistinguishable from spread noise on quiet instruments. Requiring a normalised distance beyond it selects genuine breaks and reduces the sample, which is the usual trade.
The failure condition is where definitions diverge most. Requiring a close back inside the prior range within a fixed number of bars is the strictest and most defensible. Accepting any return inside at any later point makes the rule nearly unfalsifiable, since price returns to most levels eventually and the statistic degenerates.
The time limit is not optional. Without it there is no losing case, and a rule with no losing case is not being tested.
Why does the denominator decide the result?
The result depends almost entirely on whether the sample includes the breaks that kept running, because those are the losses and they are the easiest events to leave out.
The natural way to study this pattern by eye is to find failed breakouts on a chart and examine what followed. That procedure selects on the outcome: you are looking at breaks that failed, so the finding that they reversed is built into the selection. The breaks that continued are not in the sample because they do not look like the pattern.
Mechanical enumeration fixes this by construction. Walk the history, record every break of the defined extreme that met the threshold, and grade each one against the failure condition at the stated horizon: reversed, continued, or unresolved at expiry. Report all three counts.
The ratio between those three is the actual finding, and it is the number that almost never appears in material describing the pattern. A fade rule can be perfectly real and still lose money if the continuations run far enough, which is why the size of the adverse cases matters as much as their frequency.
What makes Turtle Soup difficult to test accurately?
The specific difficulties are stop placement inside the same candle as entry, instrument-dependent noise around the extreme, and the interaction between the pattern and the prevailing regime.
Stop placement is the acute one. The stop sits just beyond the extreme that price has only just made, which puts entry and stop within a very small distance of each other, often inside a single candle. Whichever assumption the tester makes about the sequence inside that candle will shape the result, and the difference between the optimistic and honest assumptions on rules of this shape is large.
Instrument noise decides whether small breaks are real. On a pair with a wide spread relative to its range, a break of a few points beyond an extreme is inside the noise band and carries no information. Normalising the threshold by volatility handles this, and a fixed threshold does not.
Regime interaction is the analytical difficulty. A fade rule is a bet against continuation, so its results will differ sharply between trending and ranging periods. A single pooled number across a decade averages two behaviours and describes neither, and splitting the result is what turns it into something usable.
How does QuantParadox grade a rule like Turtle Soup?
QuantParadox enumerates every qualifying event mechanically across a decade of minute-resolution history on thirty instruments, resolving the entry-and-stop candle with finer data rather than an assumption.
That intrabar resolution is the decisive property for this rule specifically, because the stop is close by construction. Where minute data shows which level was reached first, that is used; where no finer resolution exists, the loss is booked rather than the win, which is the conservative direction and the only one that does not flatter the strategy.
The Reconciliation module addresses the regime question directly by identifying which conditions carry a strategy and which quietly bleed it, rather than reporting one pooled figure. For a fade rule that split is not a refinement, it is the finding.
Out-of-sample grading is applied by default, which matters because the lookback and threshold are parameters and a version tuned on one decade needs to survive a period it never saw before it means anything.