A continuously updated record rather than a selected one. That means the periods a model did badly stay on the page, because a track record that only contains good months is marketing.
Predicted vs realized
Across everything a model called at a given probability, how often did it happen?
Graded recommendations
How many outcomes stand behind each figure, so you can see when a number is thin.
Model and version
Which model produced which result. FTO-1 and FTO-1.5 are reported separately, never blended.
Period and conditions
The window each result covers and the market environment it covers, because both change what it means.
The record
Walk-forward validation · snapshot 2026-08-14
FTO-1 is reported on its 12-month walk-forward validation (1,207,199 contracts), while FTO-1.5+ adds its adaptive-exit result on the same frame and keeps its separate timing research. These are historical modeled outcomes, not live trading, and they exclude transaction costs, slippage and portfolio effects.
2.9%
FTO-1 average calibration error on displayed claims
553,000+ displayed claims; every displayed band landed within 7.3 points of what happened. Where no honest number exists, the model shows a probability range instead: 44.2% of contracts get an exact number.
0.737
FTO-1 ranking quality (ROC-AUC)
Up from 0.688 for the previous model, measured candidate-level across 1,207,199 walk-forward contracts. A ranking measure, not an accuracy percentage.
+0.0215R
FTO-1.5 modeled average per trade with adaptive exits
Against +0.0107R for the previous exit model and −0.0829R for fixed exits on the same trades. Frictionless modeled outcomes, regime-carried, not a profitability claim. Shadow evaluation.
How a probability gets made, and how it gets checked.
Outcomes, not opinions
FTO learns from a large history of trade outcomes that have each been graded against what happened. It is not encoding a view about what should work.
Graded before it counts
An outcome only enters the dataset once it can be scored. That is what makes later comparison against predicted probability meaningful rather than circular.
Evaluated after deployment
Testing a model before release tells you about the past. We keep measuring after release, because that is when a model meets conditions nobody trained it on.
Versioned, never silently swapped
When a model changes, it gets a new version and its own line in the record. Old results are not retro-fitted to new models.
Calibration
Calibration asks whether the numbers mean what they say. Across everything a model called 60%, did about 60% happen? A model can look accurate and still be badly calibrated — and a badly calibrated probability is useless for sizing a decision, which is the only thing a probability is for.
This is why we report calibration rather than a headline accuracy figure. Accuracy can be gamed by only being confident when it is easy. Calibration cannot.
Edge is not permanent
Our own research found that a model’s edge moves as conditions change — it rises, falls, recovers and drifts, and it is not a fixed property of a trained model. So we treat performance as something to track rather than a claim to make once.