Read the record before the take. A useful NFL record names wins, losses, pushes, market, date window, sample, and grading line. A percentage comes after those fields. This sounds fussy until two claims show the same rounded percentage while one rests on a large sample and the other on a small one.
The frozen NFL snapshot gives us a clean practice set: 7,548 game rows from 1999 through 2026, including 6,967 played regular-season games from 1999 through 2025. It contains closing spreads and totals. It contains no preseason rows and no opening lines. Those facts determine what records we can show and what records must remain unavailable.
Start with W-L-P
W-L-P means wins, losses, and pushes. For an ATS record, a win means the selected side beat the grading spread. A loss means it did not. A push means the final margin landed exactly on the spread. The stake is ordinarily returned on a push, but the grading record still needs to remember that the game existed.
In this snapshot, regular-season favorites from 1999 through 2025 went 3,298-3,450-189 ATS. The cover rate is 3,298 divided by 3,298 plus 3,450, or 48.9%. The 189 pushes stay visible in the record and leave the percentage denominator.
That record is more informative than “favorites covered 48.9%.” It tells you there were 6,748 decisions and 189 refunds. It also tells you the side lost more decisions than it won. The percentage is a compressed view of the line, not a replacement for it.
Name the market
A straight-up record, an ATS record, and an over-under record are different ledgers. A team can win the game and fail to cover. A game can go over while the favorite fails. Mixing markets produces a percentage that no ticket actually graded.
The same historical rows make this visible. Non-neutral home teams went 3,286-3,426-189 ATS against the closing spread. Overs went 3,399-3,469-99 against the closing total. Those records are close in percentage—49.0% and 49.5%—but they grade different claims and use different formulas.
Current picks should therefore carry an explicit market label: spread, total, moneyline, or another named market. A bare record is incomplete. Wins and losses in what market?
Name the date window
The start and end dates are part of the result. Favorites covered 53.1% in 2023 and 53.4% in 2024, then 48.3% in 2025. “Favorites have covered 53%” could be true for one selected window and false for the next completed season.
A window should be chosen before the result is used as evidence. Exploring many starts is fine if the output is labeled exploration. Publishing only the best start date as though it were the original question is not. The Analytics surface should make the filter visible so the result can be reproduced.
For a current model record, use the model version’s full graded history or a clearly named recent split. Do not silently reset the record after a weak run. Do not merge a new model into an old model’s line without a version boundary.
Name the grading line
Every ATS result in this batch uses the stored closing spread. The dictionary says positive spread_line means the home team was favored. Home ATS margin is result - spread_line. That convention is not universal inside the application, so the source and sign must be named.
An entry line and a closing line can grade the same game differently. The snapshot has only the close, so it cannot measure opening-to-closing movement or whether a recorded pick beat the close. The honest English state is “closing line available; opening line unavailable.” Our line-movement data note spells out the missing fields.
For current picks, the product should preserve the line at pick time and the declared closing line. Then it can show a plain chip: beat the close, matched the close, or missed the close. That chip is a recorded comparison, not a substitute for the W-L-P result.
Check the denominator
Small records swing. Seattle road underdogs went 10-3 ATS from 2023 through 2025, a 76.9% rate. Arizona had more covers at 13-10, a 56.5% rate. The first percentage is higher; the second record has ten more decisions. “Best” changes depending on whether the question is rate, covers, or sample.
The same problem appears in time splits. January regular-season totals produced 264 non-push decisions in the 1999–2025 sample. December produced 1,830. Printing both percentages to one decimal place does not make the evidence equally precise.
A useful card therefore shows the count beside the rate. The same rounded percentage over a small sample and a large one is not the same claim. Neither should be compressed into “wins most of the time” without the record.
Keep unavailable data unavailable
The snapshot contains 272 rows for the 2026 season and none has a final result. Those rows are scheduled games, not a graded model record. A loader should return a typed reason such as “game not played” or “result unavailable.” The UI should say what failed.
The snapshot also contains zero preseason rows. That means this article cannot report how preseason favorites, dogs, or totals performed. It also cannot generate real preseason picks from these CSVs. A blank table labeled “no results” would be better than a synthetic schedule, but an explicit sentence is better still: the data does not contain that market period.
This is the same rule for errors. A failed record query is not an empty record. “Could not load graded picks” must not render as a zero-result record. One is a system state. The other is a performance claim.
Read the price before profit
W-L-P records are comparable because they grade side outcomes. Profit is not comparable unless the price and staking rule are preserved. The CSV contains spread and total prices in many rows, but this batch does not assume every side was offered at one standard price.
A unit record should use the actual recorded price for each pick and a declared unit size. It should not backfill a convenient price onto decades of covers. A model can have a positive ATS record and a different unit result depending on prices. Both should be shown when both are real.
For a reader, the order is record, price, units. A profit percentage without W-L-P and sample can hide one large stake or an inconsistent sizing rule.
The Receipts Drawer
Here is the preseason checklist I would put beside every current model:
- Record: W-L-P, not percentage alone.
- Market: spread, total, moneyline, or a named prop.
- Window: exact start and end, plus model version.
- Sample: decisions and pushes.
- Line: entry source, grading source, and sign convention.
- Price: actual odds before any unit or return claim.
- Availability: explicit reason when a row is unplayed, missing, or failed.
The NFL picks page is where current records belong. The ATS guide covers the grading mechanics. This primer is the audit order: record first, then the story.
Three historical baselines, clearly labeled
| Historical side | Closing-line record | Rate |
|---|---|---|
| Favorites ATS | 3,298-3,450-189 | 48.9% |
| Non-neutral home teams ATS | 3,286-3,426-189 | 49.0% |
| Overs | 3,399-3,469-99 | 49.5% |
These are market baselines, not Shark Snip model records. Favorites and overs use all eligible regular-season games in the snapshot; the home-team baseline requires location = Home and excludes 66 neutral-site rows. A model record should be compared with a relevant baseline only after its own pick inclusion rule is fixed.
The table also demonstrates why the label matters. Favorites and home teams share the same 189 pushes but not the same wins and losses. Many home teams were underdogs, and some away teams were favorites. “Home favorite” is a third, narrower set.
What to do before Week 1
Do not demand a 2026 result before a 2026 game is graded. Demand a reproducible process. The pick should have a timestamp, market, side, line, price, model version, and later a final grade. The record should append; it should not be rewritten.
Historical angles can help frame questions. Week 1 dogs were 218-196-10 in the completed sample. Divisional dogs were 1,359-1,252-74. Those lines can become declared features or benchmark splits. They do not become current picks until a current model and current price select a side.
Age and risk limits still apply. Shark Snip is for adults 21 and older. Decision support is not certainty, and no historical record removes the risk of loss.
Where these numbers come from
How we counted: We used the four CSV shards under data/training-snapshots/nfl/ and the field definitions in DICTIONARY.json. The manifest records 7,548 total rows, seasons 1999–2026, and shard hashes. Completed regular-season baselines keep seasons 1999–2025 with non-null results and closing lines. Favorite ATS excludes 30 pick’em games; home ATS requires location = Home and uses result - spread_line; over-under uses total - total_line. Pushes stay in W-L-P and leave rate denominators. The 272 season-2026 rows have no result, and no row has a preseason game type, so neither group is graded.
Model calibration: predicted vs observed
Predicted win probability bucket vs the empirical win rate inside that bucket on the test set. Points on the y=x reference line are perfectly calibrated; points below mean the model is overconfident in that bucket.
NFL ATS cover-margin distribution
Distribution of (final margin − closing spread) across an NFL season. Roughly normal with mean ≈ 0 and standard deviation ≈ 13 points, which is why most ATS edges live in the ±1.5 point window.


