FTO-1 is live in the Optimizer for every user. See what changed

Don’t trust the model.

Measure it.
Advanced AI shouldn’t require blind trust. This page exists so you can decide whether ours is worth any.
Try DeriveAI free
The record
Temporal holdout · snapshot 2026-08-11
Both models are evaluated on dates they never saw in training, and reported separately rather than blended. These are temporal-holdout results, not live trading, and they exclude transaction costs, slippage and portfolio effects.
27.16%
FTO-1 top-pick target-hit rate
Against 24.28% for the strongest legacy ranker, across 75,417 held-out jobs and 325,463 candidate outcomes.
0.0229
FTO-1 calibration error (ECE, 15 bins)
Calibrated test ROC-AUC 0.7965, average precision 0.5747.
69.4%
FTO-1.5 top-5% hit rate, fixed policy
+0.541R mean realized R, versus FTO-1’s 64.5% and +0.468R. Shadow evaluation; diagnostic, not a forward result.
Calibration
Calibration asks whether the numbers mean what they say. Across everything a model called 60%, did about 60% happen? A model can look accurate and still be badly calibrated — and a badly calibrated probability is useless for sizing a decision, which is the only thing a probability is for.
This is why we report calibration rather than a headline accuracy figure. Accuracy can be gamed by only being confident when it is easy. Calibration cannot.
Edge is not permanent
Our own research found that a model’s edge moves as conditions change — it rises, falls, recovers and drifts, and it is not a fixed property of a trained model. So we treat performance as something to track rather than a claim to make once.
Read the evidence report
DeriveAI Research
Model reports and what we are learning about AI under uncertainty.
Research updates only. Unsubscribe anytime.
© Copyright DeriveAI
DeriveAI is a ROUTUR, Inc. company.
The models are built and evaluated by DeriveDX Quantitative Labs, ROUTUR's research arm.
We use analytics and session recording to see how the site is used. Form fields are never recorded. Privacy policy