Updated Sep 9, 2026 · Week 1 board.
A model that says 60% and wins 60 times in 100 is calibrated. One that says 85% and wins half its bets will cost you more than a model with a worse hit rate. Studio grades this with a single number and fails a model above 0.15.
What calibration error measures
Take every pick a model made, sort them by the confidence it claimed, then compare each claim against what happened. A model claiming 60% across a group of 100 games and winning 60 of them has no gap at all. The size of that gap, averaged across the groups, is what the check is looking at.
The check does not care whether the model wins. A model saying 40% on every pick and winning 40% of them is perfectly calibrated. That model is still useless for betting, because it never claims an edge. Calibration tests honesty, and hit rate tests skill.
Studio prints one number for it. Under 0.05 is green. Anything over 0.15 is red, and the report’s reason is that the stated probabilities do not mean what they look like.
The same check done by hand
Here is a model with 300 graded picks, sorted into four confidence groups.
| Claimed | Picks | Won | Actually won | Gap |
|---|---|---|---|---|
| 55% | 120 | 66 | 55.0% | 0 points |
| 65% | 90 | 54 | 60.0% | 5 points |
| 75% | 60 | 33 | 55.0% | 20 points |
| 85% | 30 | 15 | 50.0% | 35 points |
The quietest group is honest. The gap opens as the model gets louder, and by the top group it is claiming 85% on what turned out to be coin flips. Weighted by how many picks sit in each group, the average gap is 9 points.
Nine points is a lot. Studio computes its own version of this number in its own way, so the two are not interchangeable, and its lines sit at 0.05 and 0.15. What carries across is the shape of the failure, because this model is reliable when it is quiet and useless when it is confident.
Sorting your own picks into groups
You can run this check by hand on any model with a graded history. Sort every graded pick by the confidence the model printed on it, cut the sorted list into four or five groups, then count the wins inside each one. The claim for a group is the average confidence in it, and the truth is the share that won.
Groups need to be big enough to mean something. Thirty picks is a thin read. That makes the top group in the table above the least reliable line in it, even though it is the loudest. That is the same sample problem the trust report puts a red line under at 200.
For a model you did not build, the public track record prints calibration beside the cover rate over a grading window. Reading someone else’s groups is the cheapest way to see what a well-behaved set looks like.
Which direction the gap runs
A calibration gap has a direction, and the direction tells you what to change. The table above runs one way, because every group claims more than it delivers and the overstatement grows with the claim. That shape comes from a model rewarded for sounding confident while it was being fitted.
The opposite shape exists too. Claiming 55% and winning 62% leaves money on the table, since the stake asked for is smaller than the edge deserves. Both shapes can print the same summary number, which is why the groups are worth reading rather than the total.
The three options in the Calibration slot
Studio’s Calibration slot holds three options, and all three do the same job. They reshape raw model output into a probability you can size a bet from. Platt scaling fits a logistic curve to the claims. Isotonic calibration fits a curve that only ever climbs, and Beta calibration uses a beta-distribution fit instead.
Changing this slot and training again is the whole repair for a calibration gap. Nothing upstream fixes it by accident.
Where a gap comes from
Two things upstream push a model toward overconfidence. Too many inputs for the number of games it trained on is the common one, because the fitting starts explaining particular games instead of patterns. A training set leaning hard on one stretch of a season does the same job.
The fitting method itself is the other. Some methods hand back scores that were never probabilities to begin with, and the Calibration slot exists to repair exactly that. Studio describes the slot as the step that reshapes raw model output into a probability you can use. Fixing the gap after the fitting is cheaper than fixing the fitting, which is why the slot sits where it does.
Why this beats hit rate when money moves
Sizing reads the calibrated probability, so a wrong probability produces a wrong stake by construction. Two of the three sizing options work that way. Flat stakes is the exception, and that exception is the argument for using it while a model is young.
Take the loudest group in the table above. It claims 85%, which at even money computes a full Kelly stake of 70% of a bankroll, so the quarter-Kelly option arrives at 17.5% before its own cap is applied. Those 30 picks actually landed at 50%, which is a coin flip carrying a stake sized for near certainty.
The quieter error costs less by the same mechanism. A true 55% edge at even money computes a quarter stake of 2.5%. Claim 65% on that same bet and the quarter stake computes to 7.5%, which is three times the risk the quarter fraction was chosen to keep you inside.
The Receipts Drawer
The desk cares about this line more than it cares about hit rate. A model at 53% that knows it is at 53% is safe to stake. One that hits 60% while believing it hits 80% will find a way to lose money on a winning edge.
Calibration is also the one check that will not repair itself. Sample size improves on its own as picks grade. A calibration gap sits exactly where it is until the slot changes and the model trains again. Watch it after every retrain, because a change anywhere upstream moves it.
FAQ
What does calibration error measure? The distance between what a model claims and what happens. Sort its picks by claimed confidence, compare each group against its real hit rate, and the average gap is the idea.
My model wins 62% of its picks. Do I still need this? Yes. Two of the three sizing options read the claimed probability to decide a stake, so a dishonest number costs money on a winning model.
Why does an honest 55% model beat a 65% model that is lying? Because the stake follows the probability. At even money a true 55% edge computes a quarter-Kelly stake of 2.5% of a bankroll, while a claimed 65% computes 7.5% for the same real edge, both before the option applies its cap.
Does flat staking make calibration irrelevant? It removes the sizing damage and leaves the honesty problem. A flat staker still has to choose which picks to make, and a model that cannot rank its own confidence chooses badly at the top of its list.



