Skip to content

Benchmarks

How models compare on backtest accuracy.

Published models use the same scoring contract: Brier score, hit rate, calibration, against-the-spread rate, and drawdown. These rows are pulled from the live model registry and graded prediction results.

Registry-backed leaderboard

Last generated: Aug 28, 2026, 5:04 AM

Sport

No graded model results in the last 365 days. Metrics populate as games settle.

Benchmark rows appear after published models have enough graded predictions for the selected sport.

Metric glossary

Brier score
Mean squared error of probabilistic forecasts. Lower is better.
Log-loss
Penalizes overconfident wrong answers harder than Brier. Used for binary classifiers.
Hit rate
Win percentage excluding pushes. Breakeven against -110 juice is about 52.4%.
ATS
Against-the-spread win rate for scored spread picks.
CLV
Closing-line value: average edge captured against the final market price.
Max drawdown
Worst peak-to-trough loss in units across the scored run.

What gets audited next

  • Replayable runs: every benchmark row links back to immutable model versions and scored prediction windows.
  • Leakage checks: validation kits runs that train on the same rows they claim to evaluate.
  • Verified badges: signed benchmark runs can be surfaced on model cards and leaderboard rows.

We use cookies for essential site functionality. With your consent, we also use cookies for analytics and performance monitoring. See our Privacy Policy.