Updated Sep 9, 2026 · Studio and the desk tools.
A trained model hands you a call, not a bet. Studio's eighth slot, Sizing, holds three options that turn one into the other. Two public surfaces then decide whether the bet was any good.
Three ways to turn a call into a stake
| Sizing option | How it sets the stake | Where it fits |
|---|---|---|
| Flat stake sizing | The same fixed stake on every bet, whatever the edge | A first model, and any honest comparison between models |
| Confidence tier sizing | A stake per tier, taken from the calibrated win probability | A model whose calibration check has gone green |
| Fractional Kelly sizing | One quarter of the full Kelly stake, with a cap | A model with a long graded record behind it |
Flat sizing is the honest default while you are learning. Every bet the same size means the record you build measures the model, and not your nerve on the nights you liked one call more than the others.
Sizing is also the one slot you can change without retraining anything. The call stays exactly where it was, and only the stake attached to it moves. That makes it the cheapest experiment on the whole board.
The other two lean on the model being right about how sure it is. Both read a probability, so both inherit whatever your calibration check says about that probability. A model that says 70% and hits 55% will size its worst calls the largest.
Closing-line value is the fastest honest feedback
Graded results are slow. A full NFL season gives you a few hundred decisions, and a few hundred decisions hide a real edge inside ordinary noise.
Closing-line value asks a narrower question that answers quickly. It compares the price you took against the price the market closed at, and it does that on every bet, win or lose.
The desk publishes a free calculator for it, with no signup. Two readings of the same move come out of it. Cents movement subtracts one American price from the other, and probability delta subtracts the implied chance at your price from the implied chance at the close.
Both readings move the same way on the same bet. Cents are easier to read off a slip, and the probability version compares better across long prices and short ones. Pick one and stay with it.
Take a bet placed at plus 120 that closes at plus 105. The implied chance moves up by 3.33 percentage points, and the cents reading is 15. The closing-line value calculator prints both, alongside presets for a beaten close, a fast move, a push and a lost line.
What the public record will and will not show
The track record grades in public and holds a hard line about what counts. Only live pre-kickoff calls make the table, backfilled results cannot touch it, and losses stay on the board.
Nothing appears there until something grades. The leaderboard ranks a model only once at least 20 of its picks have been scored, and while no live pick has graded, the live section says so in plain words instead of showing a backtest dressed as a result.
The page also splits live results from backtests and labels which is which. A history that exists only as a backtest is shown as exactly that, with its sample count printed beside it, so a training record can never pass itself off as a live one.
Below that sits a calibration chart, which asks whether stated win probabilities were honest. A build claiming 80% that lands nearer 60% shows up there long before it shows up in the record.
That is the honest-absence rule, and it is the reason to trust the page. A surface that prints an empty record when the record is empty is a surface that will print your losing weeks too.
Two different twenties
Studio's publish gate wants 20 graded forward picks before it will show a forward record at all. The public leaderboard wants at least 20 scored picks before it ranks a model.
They are separate rules on separate surfaces. Clearing one does not clear the other, and a model can sit past the first and short of the second for weeks.
What the desk refuses to print
No return figure, no unit count, no dollar promise. The public page publishes cover rate, calibration and two scoring measures over a window you set to 30, 90 or 365 days.
Cover rate leads, and the sample count is printed beside it every time. That pairing is the whole design. A percentage without its sample is a claim, and a percentage with its sample is evidence.
The same restraint runs through the sizing slot. Kelly sizing is offered at a quarter stake with a cap rather than in full, because the full version maximizes growth and hands you the drawdowns to match.
The Receipts Drawer
Sizing is the slot people skip, and it is the slot that decides how a bad month feels. A model wrong about its own confidence does the most damage when the stake follows that confidence.
Closing-line value is the number to check weekly. Results need hundreds of decisions before they mean much, and every bet produces a closing price within hours.
Set flat stakes, take the closing-line reading on every bet, and leave the record alone until 20 picks have graded. If the closing readings stay positive while the record lags, keep going. Negative readings across a month point at the model, not at bad luck.
FAQ
What are the three sizing options in Studio? Flat stake sizing puts the same fixed stake on every bet whatever the edge. Confidence tier sizing sets a stake per tier, taken from the calibrated win probability. Fractional Kelly sizing stakes one quarter of the full Kelly amount, with a cap.
Which sizing option should a first model use? Flat stake sizing. It keeps every bet the same size, so the record you build measures the model rather than your staking. Move to a probability-driven option once the calibration check on the trust card is green.
What does closing-line value actually measure? The gap between the price you took and the price the market closed at. Cents movement subtracts one American price from the other. Probability delta subtracts the implied chance at your price from the implied chance at the close.
Why does the public record show nothing for my model? Because nothing has graded yet. The record page counts live pre-kickoff calls only, backfilled results cannot touch it, and the leaderboard ranks a model only once at least 20 of its picks have been scored.
How long before my model appears with a forward record? The publish gate releases a forward record once 20 forward picks have graded. That is a separate rule from the leaderboard's own 20-pick minimum, and clearing one does not clear the other.



