Skip to content
CURRENT
-7.5 → -6 (+1.5) over 24 captures ATL @ GB spread -7.5 → -6 house House backtest last 10: 6–4 · 90-day all-market 55.3% (n=8163) · 22h ago Wire Peyton fumes over INT while 'Breaking Bad' star Bryan Cranston celebrates Wire Sources: Colts' Pierce out weeks; hoping for return midseason Wire Giants QB Jaxson Dart exits MNF game vs. Rams with knee injury
Access: Anonymous access. Content follows.
learn

Why a Backtest Lies About Your Model

Read the price, role, and market first A 55% record over 60 games happens by luck about one run in four. Here is how many graded picks a record needs before it means anything at all.

9 sections

Model Desk

The Shark Snip desk for model coverage. Every claim ships with its sample size and its interval, or it does not ship.

Key takeaways (from article sections)

  • How many games a record needs
  • Break-even sits above half
  • What a backtest cannot see
  • More inputs make a backtest better and a model worse
  • Why week 1 is the worst possible test
  • The gate that ships with the model
  • The Receipts Drawer
  • FAQ
  • Where these numbers come from

Updated Sep 9, 2026 · Week 1 board.

Your first trained model will show a record, and that record is mostly noise. A fair coin flipped 60 times lands on 55% heads or better about one run in four. Nothing in that result knows anything about football.

How many games a record needs

The arithmetic here is fixed and unkind. Flip a fair coin 60 times and 55% or better arrives with probability 0.2595. Stretch the same test to 200 flips and 55% falls to 0.0895, roughly one run in eleven. At 1,000 flips it drops to 0.00087.

Graded picksA 55% record by luck aloneRead it as
60About 1 run in 4Noise
200About 1 run in 11Directional
1,000About 1 run in 1,156Evidence

Studio’s trust report draws its red line at 200 graded picks and holds green back until 1,000. Those two numbers are the same arithmetic, wearing a product label.

The same arithmetic cuts both ways. A genuinely good model can look ordinary across 60 games, because the noise that flatters a coin also buries a real edge underneath it. A small sample takes nobody’s side. It refuses to answer.

Break-even sits above half

A spread bet priced at -110 needs 52.4% to break even, because 110 divided by 210 is 0.524. Fifty percent is not the bar. The trust report’s own lines for a rolling four-week record sit either side of that figure, red below 50% and green above 52%.

Notice where green starts. A model running at 52.1% clears the report’s green line and still loses money at that price. Green on that check is a floor rather than a profit, which is the sort of thing worth knowing before you read a card as a result.

Put the two demands together and a first season gets uncomfortable. A model has to clear 52.4%, and it has to clear it across enough games that the result was not luck. Two hundred graded picks is the first point where those two demands meet each other.

What a backtest cannot see

A backtest scores a model on the games it was fitted to. Hand it enough inputs and it will describe those games beautifully, then fall apart on the next one. The trust report has a check for exactly this, and it measures how far the record moves when the training weeks change.

Studio’s Eval slot decides how the grading happens, and it holds two options. One steps through the season in order, training on the weeks before and scoring the week after, over and over. The other seals a block of games away until you publish, then grades on that block once.

Both do the same job, which is to score the model on games it has never seen. A backtest without either arrangement is a description of the past, dressed as a prediction.

More inputs make a backtest better and a model worse

Every input you add hands the fitting one more way to explain the past. Add enough of them and the backtest record climbs while the model itself gets worse, because the extra explanation is describing noise instead of football.

Studio catches both ends of that trade on the model card. Robustness measures how far the record moves when the training weeks change, which is where an overfitted model falls over. Vegas correlation catches the opposite failure, where a model has learned to copy the market and has nothing of its own left to bet.

A short backtest still has a use. Read it as a hypothesis, then ask three questions before believing it: has the sample cleared 200, did the record hold when the training weeks moved, and does the model disagree with the market at all. Two of those three sit on the card as checks in their own right, and the third is a call you make by looking at how far the model’s numbers sit from the posted line. A backtest that survives all three has earned a second look and nothing more.

Why week 1 is the worst possible test

A model trained on last season meets a league that changed over the summer. Rosters moved. Coaching staffs moved. The games that just kicked off carry no in-season history to read at all.

Studio’s Scope slot carries an option that drops season openers for this exact reason. It is worth switching on before your first training run, and it is the clearest example of a scope choice paying for itself. The longer version of the argument is in why week one breaks models.

None of that makes an opening weekend unusable. It does make an opening-weekend record worth less than the same record in November. A model whose only evidence is week 1 has been graded on the least representative games of the year.

The gate that ships with the model

Everything above is about the past. The gate below is about what happens next, and it is the only part of a model’s record that was not available to the person who built it.

Studio prints no live record for a published model until 20 graded forward picks land. The publish gate says so plainly and stays pending while the games play. A second gate waits on bets priced against a closing number before it will show any closing-line value at all.

Twenty graded picks is a small number next to the arithmetic above, and it is meant to be. That gate proves nothing on its own. It marks the point where Studio stops printing a blank and starts printing a number that came from games nobody had played when the model went out.

Two honesty laws sit above all of this, printed on the board where you build. One bars the settled result of a game from ever becoming an input. The other bars the closing line, because that number arrives after the moment the model has to speak.

The Receipts Drawer

The desk position on this is short. A backtest is a hypothesis and a forward record is the test of it, and the two deserve different levels of belief. Anyone quoting a 60-game backtest is quoting a coin.

Train the model, save it, then leave it alone. Twenty graded picks is the first honest read and 200 is the first stable one. If the number moves a long way between those two points, the model is telling you it was fitted to noise, and the repair belongs in Scope and Transforms rather than in another training run.

FAQ

How many bets do I need before a record means anything? Two hundred graded picks moves the trust report off red, and a thousand turns it green. Under 200 the report treats the record as noise, and the arithmetic agrees with it.

How does a forward record differ from a backtest? A backtest scores the model on games it was fitted to. A forward record scores it on games that had not been played when the model was published. Studio waits for 20 of those before printing one.

Why is the closing line barred as an input? A printed honesty law bars it. The closing line arrives after the market has finished pricing the game, so a model reading it answers a question it already holds.

Which do I believe, my backtest or my live record? The live one, once 20 picks have graded. Read the robustness check as well, because a record that swings with the training weeks was never stable enough to trust.

Where these numbers come from

Frequently asked questions

How many bets do I need before a record means anything?
Two hundred graded picks moves the trust report off red, and a thousand turns it green. Under 200 the report treats the record as noise, and the arithmetic agrees with it.
How does a forward record differ from a backtest?
A backtest scores the model on games it was fitted to. A forward record scores it on games that had not been played when the model was published. Studio waits for 20 of those before printing one.
Why is the closing line barred as an input?
A printed honesty law bars it. The closing line arrives after the market has finished pricing the game, so a model reading it answers a question it already holds.
Which do I believe, my backtest or my live record?
The live one, once 20 picks have graded. Read the robustness check as well, because a record that swings with the training weeks was never stable enough to trust.

Build a free model in 60 seconds →

Go →
7m read time
3 players/teams
8 key angles

Angles in this read

  • Probability bands Ranges and uncertainty are shown as bands rather than fake certainty.
  • Odds tick Micro tick movement reinforces live market and pricing language.
  • Line arrow Spread, total, and price movement sections get directional cues.
  • Edge meter Positive expected value is presented as a meter, not a guarantee.
  • Model sparkline Model output and projection movement get a tiny sparkline rhythm.
  • Research scan Tables, evidence ledgers, and inline charts receive a research-note scan cue.

This article's context stays anchored to Updated Sep, Evidence Studio and FAQ How and closing line value, model and price, all of which appear in the post itself.

Names and terms found in this article
Updated SepEvidence StudioFAQ Howclosing line valuemodelpricebuildershark snip builder
Share this guide Help another reader make a sharper decision.

Get picks in your inbox

One email, every slate — ranked edges, no touts. Unsubscribe any time.

Start free — pick a sport

Go →

Continue with evidence

Related reading and source status

Related Reads

Inside the Six-Stage Reveal When Training Finishes — on the Advanced Canvas — Shark Snip
Betting Tools

Inside the Six-Stage Reveal When Training Finishes — on the Advanced Canvas

A finished training run on /build/[slug] plays a fixed six-stage reveal, grade first, timed to the exact millisecond in the source.

Sep 11, 2026 5 min read
The Screen That Shows Your Backend Before Training Starts — on the Advanced Canvas — Shark Snip
Betting Tools

The Screen That Shows Your Backend Before Training Starts — on the Advanced Canvas

Pre-Train Preview on /build/[slug] names the exact backend a run will use, from a single side-effect-free probe, before training starts.

Sep 11, 2026 5 min read
Eight Lattice Slots, and the 56 Modules That Can Fill Them — Shark Snip
Betting Tools

Eight Lattice Slots, and the 56 Modules That Can Fill Them

The /build board holds eight fixed slots filled from a live, growing list of pieces, not a four-tier gallery.

Sep 11, 2026 6 min read

query: loadMergedBlogPostCards + scoreRelated · n = 3

No data

No graded source picks match this article yet

The public.source_accuracy_scores 90-day query returned no rows for this article's inferred sport.