Updated Sep 9, 2026 · Week 1 board.
Your first trained model will show a record, and that record is mostly noise. A fair coin flipped 60 times lands on 55% heads or better about one run in four. Nothing in that result knows anything about football.
How many games a record needs
The arithmetic here is fixed and unkind. Flip a fair coin 60 times and 55% or better arrives with probability 0.2595. Stretch the same test to 200 flips and 55% falls to 0.0895, roughly one run in eleven. At 1,000 flips it drops to 0.00087.
| Graded picks | A 55% record by luck alone | Read it as |
|---|---|---|
| 60 | About 1 run in 4 | Noise |
| 200 | About 1 run in 11 | Directional |
| 1,000 | About 1 run in 1,156 | Evidence |
Studio’s trust report draws its red line at 200 graded picks and holds green back until 1,000. Those two numbers are the same arithmetic, wearing a product label.
The same arithmetic cuts both ways. A genuinely good model can look ordinary across 60 games, because the noise that flatters a coin also buries a real edge underneath it. A small sample takes nobody’s side. It refuses to answer.
Break-even sits above half
A spread bet priced at -110 needs 52.4% to break even, because 110 divided by 210 is 0.524. Fifty percent is not the bar. The trust report’s own lines for a rolling four-week record sit either side of that figure, red below 50% and green above 52%.
Notice where green starts. A model running at 52.1% clears the report’s green line and still loses money at that price. Green on that check is a floor rather than a profit, which is the sort of thing worth knowing before you read a card as a result.
Put the two demands together and a first season gets uncomfortable. A model has to clear 52.4%, and it has to clear it across enough games that the result was not luck. Two hundred graded picks is the first point where those two demands meet each other.
What a backtest cannot see
A backtest scores a model on the games it was fitted to. Hand it enough inputs and it will describe those games beautifully, then fall apart on the next one. The trust report has a check for exactly this, and it measures how far the record moves when the training weeks change.
Studio’s Eval slot decides how the grading happens, and it holds two options. One steps through the season in order, training on the weeks before and scoring the week after, over and over. The other seals a block of games away until you publish, then grades on that block once.
Both do the same job, which is to score the model on games it has never seen. A backtest without either arrangement is a description of the past, dressed as a prediction.
More inputs make a backtest better and a model worse
Every input you add hands the fitting one more way to explain the past. Add enough of them and the backtest record climbs while the model itself gets worse, because the extra explanation is describing noise instead of football.
Studio catches both ends of that trade on the model card. Robustness measures how far the record moves when the training weeks change, which is where an overfitted model falls over. Vegas correlation catches the opposite failure, where a model has learned to copy the market and has nothing of its own left to bet.
A short backtest still has a use. Read it as a hypothesis, then ask three questions before believing it: has the sample cleared 200, did the record hold when the training weeks moved, and does the model disagree with the market at all. Two of those three sit on the card as checks in their own right, and the third is a call you make by looking at how far the model’s numbers sit from the posted line. A backtest that survives all three has earned a second look and nothing more.
Why week 1 is the worst possible test
A model trained on last season meets a league that changed over the summer. Rosters moved. Coaching staffs moved. The games that just kicked off carry no in-season history to read at all.
Studio’s Scope slot carries an option that drops season openers for this exact reason. It is worth switching on before your first training run, and it is the clearest example of a scope choice paying for itself. The longer version of the argument is in why week one breaks models.
None of that makes an opening weekend unusable. It does make an opening-weekend record worth less than the same record in November. A model whose only evidence is week 1 has been graded on the least representative games of the year.
The gate that ships with the model
Everything above is about the past. The gate below is about what happens next, and it is the only part of a model’s record that was not available to the person who built it.
Studio prints no live record for a published model until 20 graded forward picks land. The publish gate says so plainly and stays pending while the games play. A second gate waits on bets priced against a closing number before it will show any closing-line value at all.
Twenty graded picks is a small number next to the arithmetic above, and it is meant to be. That gate proves nothing on its own. It marks the point where Studio stops printing a blank and starts printing a number that came from games nobody had played when the model went out.
Two honesty laws sit above all of this, printed on the board where you build. One bars the settled result of a game from ever becoming an input. The other bars the closing line, because that number arrives after the moment the model has to speak.
The Receipts Drawer
The desk position on this is short. A backtest is a hypothesis and a forward record is the test of it, and the two deserve different levels of belief. Anyone quoting a 60-game backtest is quoting a coin.
Train the model, save it, then leave it alone. Twenty graded picks is the first honest read and 200 is the first stable one. If the number moves a long way between those two points, the model is telling you it was fitted to noise, and the repair belongs in Scope and Transforms rather than in another training run.
FAQ
How many bets do I need before a record means anything? Two hundred graded picks moves the trust report off red, and a thousand turns it green. Under 200 the report treats the record as noise, and the arithmetic agrees with it.
How does a forward record differ from a backtest? A backtest scores the model on games it was fitted to. A forward record scores it on games that had not been played when the model was published. Studio waits for 20 of those before printing one.
Why is the closing line barred as an input? A printed honesty law bars it. The closing line arrives after the market has finished pricing the game, so a model reading it answers a question it already holds.
Which do I believe, my backtest or my live record? The live one, once 20 picks have graded. Read the robustness check as well, because a record that swings with the training weeks was never stable enough to trust.



