A continuously updated record rather than a selected one. That means the periods a model did badly stay on the page, because a track record that only contains good months is marketing.
Predicted vs realized
Across everything a model called at a given probability, how often did it happen?
Graded recommendations
How many outcomes stand behind each figure, so you can see when a number is thin.
Model and version
Which model produced which result. FTO-1 and FTO-1.5 are reported separately, never blended.
Period and conditions
The window each result covers and the market environment it covers, because both change what it means.
The record
Temporal holdout · snapshot 2026-08-11
Both models are evaluated on dates they never saw in training, and reported separately rather than blended. These are temporal-holdout results, not live trading, and they exclude transaction costs, slippage and portfolio effects.
27.16%
FTO-1 top-pick target-hit rate
Against 24.28% for the strongest legacy ranker, across 75,417 held-out jobs and 325,463 candidate outcomes.
0.0229
FTO-1 calibration error (ECE, 15 bins)
Calibrated test ROC-AUC 0.7965, average precision 0.5747.
69.4%
FTO-1.5 top-5% hit rate, fixed policy
+0.541R mean realized R, versus FTO-1’s 64.5% and +0.468R. Shadow evaluation; diagnostic, not a forward result.
How a probability gets made, and how it gets checked.
Outcomes, not opinions
FTO learns from a large history of trade outcomes that have each been graded against what happened. It is not encoding a view about what should work.
Graded before it counts
An outcome only enters the dataset once it can be scored. That is what makes later comparison against predicted probability meaningful rather than circular.
Evaluated after deployment
Testing a model before release tells you about the past. We keep measuring after release, because that is when a model meets conditions nobody trained it on.
Versioned, never silently swapped
When a model changes, it gets a new version and its own line in the record. Old results are not retro-fitted to new models.
Calibration
Calibration asks whether the numbers mean what they say. Across everything a model called 60%, did about 60% happen? A model can look accurate and still be badly calibrated — and a badly calibrated probability is useless for sizing a decision, which is the only thing a probability is for.
This is why we report calibration rather than a headline accuracy figure. Accuracy can be gamed by only being confident when it is easy. Calibration cannot.
Edge is not permanent
Our own research found that a model’s edge moves as conditions change — it rises, falls, recovers and drifts, and it is not a fixed property of a trained model. So we treat performance as something to track rather than a claim to make once.