Skip to content
CURRENT
-7.5 → -6 (+1.5) over 24 captures ATL @ GB spread -7.5 → -6 house House backtest last 10: 6–4 · 90-day all-market 55.3% (n=8163) · 22h ago Wire Peyton fumes over INT while 'Breaking Bad' star Bryan Cranston celebrates Wire Sources: Colts' Pierce out weeks; hoping for return midseason Wire Giants QB Jaxson Dart exits MNF game vs. Rams with knee injury
Access: Anonymous access. Content follows.
learn

Calibration: When 60% Means 60%

Read the price, role, and market first A model saying 60% should win about 60 times in 100. Here is how to check that by hand, what Studio measures, and why it changes the size of a bet.

10 sections

Model Desk

The Shark Snip desk for model coverage. Every claim ships with its sample size and its interval, or it does not ship.

Key takeaways (from article sections)

  • What calibration error measures
  • The same check done by hand
  • Sorting your own picks into groups
  • Which direction the gap runs
  • The three options in the Calibration slot
  • Where a gap comes from
  • Why this beats hit rate when money moves
  • The Receipts Drawer
  • FAQ
  • Where these numbers come from

Updated Sep 9, 2026 · Week 1 board.

A model that says 60% and wins 60 times in 100 is calibrated. One that says 85% and wins half its bets will cost you more than a model with a worse hit rate. Studio grades this with a single number and fails a model above 0.15.

What calibration error measures

Take every pick a model made, sort them by the confidence it claimed, then compare each claim against what happened. A model claiming 60% across a group of 100 games and winning 60 of them has no gap at all. The size of that gap, averaged across the groups, is what the check is looking at.

The check does not care whether the model wins. A model saying 40% on every pick and winning 40% of them is perfectly calibrated. That model is still useless for betting, because it never claims an edge. Calibration tests honesty, and hit rate tests skill.

Studio prints one number for it. Under 0.05 is green. Anything over 0.15 is red, and the report’s reason is that the stated probabilities do not mean what they look like.

The same check done by hand

Here is a model with 300 graded picks, sorted into four confidence groups.

ClaimedPicksWonActually wonGap
55%1206655.0%0 points
65%905460.0%5 points
75%603355.0%20 points
85%301550.0%35 points

The quietest group is honest. The gap opens as the model gets louder, and by the top group it is claiming 85% on what turned out to be coin flips. Weighted by how many picks sit in each group, the average gap is 9 points.

Nine points is a lot. Studio computes its own version of this number in its own way, so the two are not interchangeable, and its lines sit at 0.05 and 0.15. What carries across is the shape of the failure, because this model is reliable when it is quiet and useless when it is confident.

Sorting your own picks into groups

You can run this check by hand on any model with a graded history. Sort every graded pick by the confidence the model printed on it, cut the sorted list into four or five groups, then count the wins inside each one. The claim for a group is the average confidence in it, and the truth is the share that won.

Groups need to be big enough to mean something. Thirty picks is a thin read. That makes the top group in the table above the least reliable line in it, even though it is the loudest. That is the same sample problem the trust report puts a red line under at 200.

For a model you did not build, the public track record prints calibration beside the cover rate over a grading window. Reading someone else’s groups is the cheapest way to see what a well-behaved set looks like.

Which direction the gap runs

A calibration gap has a direction, and the direction tells you what to change. The table above runs one way, because every group claims more than it delivers and the overstatement grows with the claim. That shape comes from a model rewarded for sounding confident while it was being fitted.

The opposite shape exists too. Claiming 55% and winning 62% leaves money on the table, since the stake asked for is smaller than the edge deserves. Both shapes can print the same summary number, which is why the groups are worth reading rather than the total.

The three options in the Calibration slot

Studio’s Calibration slot holds three options, and all three do the same job. They reshape raw model output into a probability you can size a bet from. Platt scaling fits a logistic curve to the claims. Isotonic calibration fits a curve that only ever climbs, and Beta calibration uses a beta-distribution fit instead.

Changing this slot and training again is the whole repair for a calibration gap. Nothing upstream fixes it by accident.

Where a gap comes from

Two things upstream push a model toward overconfidence. Too many inputs for the number of games it trained on is the common one, because the fitting starts explaining particular games instead of patterns. A training set leaning hard on one stretch of a season does the same job.

The fitting method itself is the other. Some methods hand back scores that were never probabilities to begin with, and the Calibration slot exists to repair exactly that. Studio describes the slot as the step that reshapes raw model output into a probability you can use. Fixing the gap after the fitting is cheaper than fixing the fitting, which is why the slot sits where it does.

Why this beats hit rate when money moves

Sizing reads the calibrated probability, so a wrong probability produces a wrong stake by construction. Two of the three sizing options work that way. Flat stakes is the exception, and that exception is the argument for using it while a model is young.

Take the loudest group in the table above. It claims 85%, which at even money computes a full Kelly stake of 70% of a bankroll, so the quarter-Kelly option arrives at 17.5% before its own cap is applied. Those 30 picks actually landed at 50%, which is a coin flip carrying a stake sized for near certainty.

The quieter error costs less by the same mechanism. A true 55% edge at even money computes a quarter stake of 2.5%. Claim 65% on that same bet and the quarter stake computes to 7.5%, which is three times the risk the quarter fraction was chosen to keep you inside.

The Receipts Drawer

The desk cares about this line more than it cares about hit rate. A model at 53% that knows it is at 53% is safe to stake. One that hits 60% while believing it hits 80% will find a way to lose money on a winning edge.

Calibration is also the one check that will not repair itself. Sample size improves on its own as picks grade. A calibration gap sits exactly where it is until the slot changes and the model trains again. Watch it after every retrain, because a change anywhere upstream moves it.

FAQ

What does calibration error measure? The distance between what a model claims and what happens. Sort its picks by claimed confidence, compare each group against its real hit rate, and the average gap is the idea.

My model wins 62% of its picks. Do I still need this? Yes. Two of the three sizing options read the claimed probability to decide a stake, so a dishonest number costs money on a winning model.

Why does an honest 55% model beat a 65% model that is lying? Because the stake follows the probability. At even money a true 55% edge computes a quarter-Kelly stake of 2.5% of a bankroll, while a claimed 65% computes 7.5% for the same real edge, both before the option applies its cap.

Does flat staking make calibration irrelevant? It removes the sizing damage and leaves the honesty problem. A flat staker still has to choose which picks to make, and a model that cannot rank its own confidence chooses badly at the top of its list.

Where these numbers come from

Frequently asked questions

What does calibration error measure?
The distance between what a model claims and what happens. Sort its picks by claimed confidence, compare each group against its real hit rate, and the average gap is the idea.
My model wins 62% of its picks. Do I still need this?
Yes. Two of the three sizing options read the claimed probability to decide a stake, so a dishonest number costs money on a winning model.
Why does an honest 55% model beat a 65% model that is lying?
Because the stake follows the probability. At even money a true 55% edge computes a quarter-Kelly stake of 2.5% of a bankroll, while a claimed 65% computes 7.5% for the same real edge, both before the option applies its cap.
Does flat staking make calibration irrelevant?
It removes the sizing damage and leaves the honesty problem. A flat staker still has to choose which picks to make, and a model that cannot rank its own confidence chooses badly at the top of its list.

Build a free model in 60 seconds →

Go →
7m read time
4 players/teams
8 key angles

Angles in this read

  • Edge meter Positive expected value is presented as a meter, not a guarantee.
  • Probability bands Ranges and uncertainty are shown as bands rather than fake certainty.
  • Model sparkline Model output and projection movement get a tiny sparkline rhythm.
  • Line arrow Spread, total, and price movement sections get directional cues.
  • Research scan Tables, evidence ledgers, and inline charts receive a research-note scan cue.
  • Line reveal Pretext-measured lines reveal without reflowing the article.

This article's context stays anchored to Updated Sep, Claimed Picks Won Actually and FAQ What and model, builder and shark snip builder, all of which appear in the post itself.

Names and terms found in this article
Updated SepClaimed Picks Won ActuallyFAQ WhatYes. Twomodelbuildershark snip builderhow tobuild page
Share this guide Help another reader make a sharper decision.

Get picks in your inbox

One email, every slate — ranked edges, no touts. Unsubscribe any time.

Start free — pick a sport

Go →

Continue with evidence

Related reading and source status

Related Reads

Inside the Six-Stage Reveal When Training Finishes — on the Advanced Canvas — Shark Snip
Betting Tools

Inside the Six-Stage Reveal When Training Finishes — on the Advanced Canvas

A finished training run on /build/[slug] plays a fixed six-stage reveal, grade first, timed to the exact millisecond in the source.

Sep 11, 2026 5 min read
The Screen That Shows Your Backend Before Training Starts — on the Advanced Canvas — Shark Snip
Betting Tools

The Screen That Shows Your Backend Before Training Starts — on the Advanced Canvas

Pre-Train Preview on /build/[slug] names the exact backend a run will use, from a single side-effect-free probe, before training starts.

Sep 11, 2026 5 min read
Eight Lattice Slots, and the 56 Modules That Can Fill Them — Shark Snip
Betting Tools

Eight Lattice Slots, and the 56 Modules That Can Fill Them

The /build board holds eight fixed slots filled from a live, growing list of pieces, not a four-tier gallery.

Sep 11, 2026 6 min read

query: loadMergedBlogPostCards + scoreRelated · n = 3

No data

No graded source picks match this article yet

The public.source_accuracy_scores 90-day query returned no rows for this article's inferred sport.