| Test scope | Result | |
|---|---|---|
Ranking accuracy (ROC-AUC) How often a target hitter ranks above a non-hitter; 0.50 is chance and 1.00 is perfect. Candidate-level, walk-forward validated. | All 1,207,199 candidates · 12 monthly folds | 0.7366 |
Ranking accuracy, prior model The deployed model this release replaces, on the same test. | All 1,207,199 candidates · 12 monthly folds | 0.6880 |
Confidence penalty (log loss) Penalizes confident predictions that are wrong; lower is better. Prior model: 0.5104. | All 1,207,199 candidates | 0.4665 |
Calibration error on displayed claims Average gap between the probabilities shown and the outcomes observed; lower is better. | 553,000+ displayed claims | 2.9% |
Calibration error, unseen half-span The same measurement restricted to the second half of the validation span, which no tuning decision saw. | Displayed claims, second half | 4.2% |
Worst displayed band The largest gap between any displayed probability band and what it delivered. | 10 probability bands | 7.3pp |
Numeric coverage The share of contracts shown an exact number. The model abstains rather than display an unsupported number; the rest get a probability range. | All displayed contracts | 44.2% |
Stated ~6% band Observed barrier-hit rate against the stated probability. Intervals in parentheses are 95% confidence ranges where computed. | 250,053 displayed claims | 7.9% observed (7.8–8.0%) |
Stated ~16% band | 97,200 displayed claims | 19.9% observed |
Stated ~25% band | 90,125 displayed claims | 24.1% observed (23.8–24.4%) |
Stated ~35% band | 49,878 displayed claims | 29.9% observed |
Stated ~54% band | 8,656 displayed claims | 47.3% observed |
Stated ~64% band | 13,176 displayed claims | 60.0% observed (59.2–60.8%) |
Stated ~76% band | 12,711 displayed claims | 68.9% observed |
Stated ~94% band | 12,348 displayed claims | 88.6% observed (88.0–89.1%) |