Benchmarks
How models compare on backtest accuracy.
Published models use the same scoring contract: Brier score, hit rate, calibration, against-the-spread rate, and drawdown. These rows are pulled from the live model registry and graded prediction results.
Registry-backed leaderboard
Last generated: Aug 28, 2026, 5:04 AM
Sport
No graded model results in the last 365 days. Metrics populate as games settle.
Benchmark rows appear after published models have enough graded predictions for the selected sport.
Metric glossary
- Brier score
- Mean squared error of probabilistic forecasts. Lower is better.
- Log-loss
- Penalizes overconfident wrong answers harder than Brier. Used for binary classifiers.
- Hit rate
- Win percentage excluding pushes. Breakeven against -110 juice is about 52.4%.
- ATS
- Against-the-spread win rate for scored spread picks.
- CLV
- Closing-line value: average edge captured against the final market price.
- Max drawdown
- Worst peak-to-trough loss in units across the scored run.
What gets audited next
- Replayable runs: every benchmark row links back to immutable model versions and scored prediction windows.
- Leakage checks: validation kits runs that train on the same rows they claim to evaluate.
- Verified badges: signed benchmark runs can be surfaced on model cards and leaderboard rows.
