title: "Methodology: how our NFL model works, and how it performed against the market" description: "The full model specification, walk-forward backtest results including the ones that go against us, the calibration table, the confidence cap, the leakage controls, and how the SHA-256 Merkle commitment works." last_updated: 2026-08-02 model_version: elo-mov-v1+iso
This page is the complete technical description of how our predictions are produced and how they performed in testing. It includes the results that argue against paying us for picks. That is deliberate. It is also the file we would hand to a regulator, a payment processor, or a journalist who asked us to substantiate anything on this site.
Everything on this page is reproducible from the public ledger and the open source code. If a number here cannot be regenerated from published data, treat it as an error and tell us.
We predict NFL game outcomes with an Elo rating system that uses a damped margin-of-victory update, then convert those ratings into probabilities and correct those probabilities with isotonic regression refit each season on prior seasons only. Across 2,750 out-of-sample games from 2016 through 2025, the model recorded a Brier score of 0.22228 against the de-vigged closing market's 0.21061, straight-up accuracy of 63.96% against the market's 66.58%, and an against-the-spread record of 1333-1352-65, or 49.65%, against a breakeven of 52.38% at standard -110 pricing. The model does not beat the closing market. Our probabilities are well calibrated, which is a different and smaller claim, and it is the only claim we make.
We claim exactly two things:
We do not claim, and have never claimed, that following our predictions is profitable. Our own testing says it is not. We do not publish a win rate as a selling point, a return on investment figure, or a units-won figure. We do not accept wagers, hold funds, or pay out prizes.
Every team carries a single rating. Ratings start at 1500 and move only when a game settles.
The probability that the home team wins is the standard Elo logistic:
p_home = 1 / (1 + 10 ^ (-(R_home - R_away + HFA + rest_bonus) / 400))
After the game settles, both ratings move by the same amount in opposite directions:
delta = K * mov_multiplier * (actual - p_home)
mov_multiplier = ln(|margin| + 1) * (2.2 / (0.001 * winner_elo_diff + 2.2))
The margin-of-victory multiplier is the standard log-damped form. It gives more
credit for a 24-point win than a 3-point win without letting a single blowout
dominate a team's rating, and the winner_elo_diff denominator corrects for the
fact that better teams are mechanically more likely to win by a lot. Without
that correction the system over-rates favourites in a self-reinforcing loop.
| parameter | value | meaning |
|---|---|---|
k |
20.0 | rating points moved per unit of surprise |
home_advantage |
48.0 Elo | roughly 2 points of spread |
season_carryover |
0.75 | fraction of a rating carried into the next season |
base_rating |
1500.0 | starting and mean-reversion target |
elo_per_point |
25.0 | Elo points per point of scoring margin |
rest_per_day |
1.5 Elo | bonus per extra day of rest versus the opponent |
These values are conventional, were not tuned against the test period, and are published so that anyone can reproduce our ratings exactly. The model is deliberately simple. Its job is not to be clever; its job is to be a fully explainable baseline that we can publish, grade in public, and improve on in the open.
Raw Elo probabilities are mildly overconfident in the middle bands. We correct them with isotonic regression, which is a monotone step-fit that reshapes the probability curve without assuming that curve is logistic.
The calibrator is fitted only on seasons strictly before the season being predicted, and it is refitted every year. Fitting a calibrator on the same games you then score with it manufactures a perfect-looking reliability curve that means nothing. Seasons before we have 500 prior games available pass through uncalibrated rather than being calibrated on thin data.
The production model version is elo-mov-v1+iso.
Ratings are built strictly forward in time. For every game in the historical record the model produces a prediction using only ratings derived from games that had already finished, and only afterwards is the result used to update the ratings. There is no point at which the model sees a future game.
The market comparison uses de-vigged closing prices. Both sides' implied probabilities are divided by their sum to remove the bookmaker margin. Comparing a model to a vigged line is trivially easy and meaningless, because the vig guarantees the raw line is a biased probability estimate.
Every figure below is regenerated by a single command:
python scripts/published_figures.py
If a number on this site cannot be produced by that command, it should not be on this site. We publish two evaluations of the same models rather than choosing the flattering one.
Larger sample, weaker provenance. nflverse's spread_line is an undocumented
periodic snapshot rather than a documented close.
| model | n | Brier | ECE | ATS record | ATS% |
|---|---|---|---|---|---|
| Elo baseline | 2671 | 0.22217 | 0.02698 | 1287-1321-63 | 0.4935 |
| Independent (ours) | 2671 | 0.22148 | 0.03122 | 1291-1317-63 | 0.4950 |
| Consensus (+market) | 2671 | 0.21455 | 0.03343 | 1300-1308-63 | 0.4985 |
| Closing market | 2671 | 0.21038 | 0.01913 | 1327-1281-63 | 0.5088 |
Smaller sample, far better provenance: the median across books, captured 5–28 minutes before each kickoff, from odds we paid for and hold ourselves.
| model | n | Brier | ECE | ATS record | ATS% |
|---|---|---|---|---|---|
| Elo baseline | 854 | 0.22246 | 0.03265 | 402-431-21 | 0.4826 |
| Independent (ours) | 854 | 0.22164 | 0.03269 | 406-427-21 | 0.4874 |
| Consensus (+market) | 854 | 0.21399 | 0.04231 | 412-421-21 | 0.4946 |
| Closing market | 854 | 0.21061 | 0.02396 | 406-427-21 | 0.4874 |
At −110 on both sides a bettor needs 52.38% against the spread to break even. No model in either table clears it. Neither does the closing market against its own number — which is the sanity check that the test is well-formed rather than flattering us.
32.9% of spreads differ between nflverse and the real consensus close, but the typical difference is small: mean 0.217 points, median 0.0, and only 5.5% differ by a full point or more. Graded on the same games, the Elo baseline scores 0.4850 against nflverse lines and 0.4826 against real closes — a 0.24 percentage point gap. The weaker source was precise enough for the conclusion and not precise enough to publish closing-line value from, which is why we bought the better one.
Calibration asks a narrower question than profitability: when we say 70%, does it happen about 70% of the time? A model can be well calibrated and still unprofitable, which is precisely our situation.
Expected calibration error (ECE) is the sample-weighted mean absolute gap between predicted and observed frequency across ten probability bands. Lower is better.
| model | ECE | Brier |
|---|---|---|
| raw Elo | 0.02654 | 0.22228 |
| isotonic-calibrated Elo (published) | 0.02162 | 0.22323 |
| de-vigged market | 0.01802 | 0.21061 |
Two honest notes on that table. First, calibration improves ECE by about 18.5% but very slightly worsens Brier score, from 0.22228 to 0.22323. Isotonic regression buys reliability at a small cost in sharpness. We publish both numbers because reporting only the one that improved would be the same selective disclosure we are criticising. Second, the market is still better calibrated than we are.
| predicted band | n | mean predicted | actual frequency | gap |
|---|---|---|---|---|
| 0.0-0.1 | 12 | 1.00% | 33.33% | -32.33 pts |
| 0.1-0.2 | 18 | 16.66% | 27.78% | -11.12 pts |
| 0.2-0.3 | 149 | 24.84% | 30.87% | -6.03 pts |
| 0.3-0.4 | 324 | 36.21% | 34.88% | +1.33 pts |
| 0.4-0.5 | 593 | 44.80% | 43.00% | +1.80 pts |
| 0.5-0.6 | 401 | 56.10% | 53.62% | +2.48 pts |
| 0.6-0.7 | 739 | 64.98% | 63.87% | +1.11 pts |
| 0.7-0.8 | 244 | 73.66% | 72.54% | +1.12 pts |
| 0.8-0.9 | 210 | 84.77% | 82.86% | +1.91 pts |
| 0.9-1.0 | 60 | 94.51% | 86.67% | +7.84 pts |
A positive gap means we were overconfident: we predicted the event more often than it happened.
Read the extremes with care. The 0.0-0.1 and 0.1-0.2 bands hold 12 and 18 games respectively; at those sample sizes a handful of upsets moves the observed frequency by tens of points and the gap is mostly noise. The bands that carry real weight are 0.3 through 0.9, which hold 2,511 of the 2,750 games, and in those bands the model is overconfident by between 1.1 and 2.5 percentage points.
For comparison, the de-vigged market over the same games is mildly underconfident in its top bands: it predicted 84.70% and observed 87.57% in the 0.8-0.9 band. That is the signature of a market that has priced in the vig asymmetry, and it is another reason we treat the market as the benchmark rather than the opponent.
We do not publish any probability above 0.85, regardless of what the model outputs. The cap exists because of one row in the table above.
In the 0.9-1.0 band, across 60 out-of-sample games, the calibrated model predicted an average of 94.51% and the events occurred 86.67% of the time. The model was overconfident by 7.84 percentage points - four times the error of any other well-populated band.
This is the single most important finding in our testing, because it inverts the industry's usual sales pitch. The standard product in this category is "our one highest-confidence pick of the day." That is precisely the band where our model is least trustworthy. High-confidence selections are not the safe subset; in our data they are the least reliable subset, because extreme probabilities are produced by extreme rating gaps, and extreme rating gaps are exactly where a simple rating system is most likely to be extrapolating past the evidence.
Capping at 0.85 keeps every published probability inside a band where our measured overconfidence is under two percentage points. It costs us the headline number that would sell best. We consider explaining why we do not have that number to be more valuable than having it.
Data leakage - training on information that would not have been available before kickoff - is the most common way a sports model backtests beautifully and then loses live. It is also the easiest way to build a fraudulent-looking track record without meaning to.
The following columns are present in our source data and are banned from feature construction, enforced by an assertion in the NFL adapter that fails the build rather than warning:
| banned column | why |
|---|---|
temp |
measured at or after the game, not a pre-game forecast |
wind |
same |
home_score, away_score |
the outcome |
result |
the outcome |
total |
the outcome |
overtime |
the outcome |
Betting-line columns require a separate control. In our source dataset,
spread_line and the related price columns are overwritten in place as the
market moves. A row read today shows the current number, not the number that
existed when the game was scheduled. We therefore treat those columns as a
closing line only for games already marked final, and we never treat them as an
opening line.
Because of that, we have not published a closing-line-value figure and will not publish one until we have validated our line history against an independent source with explicit open and close objects. CLV is the metric most often quoted by services in this category and it is the metric most easily faked by reading a mutable field. Ours will be published when it is defensible and not before.
The features actually used by the production model are: team identity, prior ratings derived only from earlier games, home or neutral site, days of rest for each team, and season boundaries for the carryover regression. That is the complete list.
The problem with every published pick record on the internet is that it is self-reported and editable. A losing pick can quietly vanish. A winning pick can be added afterwards. "We publish everything" is an unfalsifiable claim.
We make our record falsifiable with a commit-reveal scheme.
leaf = SHA-256(0x00 || canonical_json)parent = SHA-256(0x01 || left || right), working up the tree until one node
remains. If a level has an odd number of nodes, the last node is duplicated.
The distinct 0x00 and 0x01 prefixes are domain separation: they make it
impossible to pass an internal node off as a leaf, which is the standard
second-preimage attack on naive Merkle trees.Anyone can then recompute the tree from the revealed predictions and confirm it
produces the root we published before kickoff. If we had altered, removed,
reordered, or back-dated a single prediction, the recomputed root would not
match. The algorithm identifier recorded in every commitment file is
sha256-merkle-v1.
The commit function refuses to seal a slate after its first kickoff. A commitment created after games have started proves nothing, so we made it impossible to create one by accident.
Because it is a Merkle tree and not a flat hash, a single prediction can be proven to belong to a committed slate without revealing the rest of the slate. For our 16-game NFL Week 1 slate, an inclusion proof is four hashes long. This matters for the paid tier: we can prove a subscriber-only prediction was part of the pre-kickoff commitment without publishing the whole slate to non-subscribers.
Step-by-step instructions for verifying all of this yourself, including a 30-line script that does not use any of our code, are on the verification page.
It proves the record is complete and unedited. Every prediction we made is in it, in the form we made it, timestamped before kickoff.
It does not prove the predictions are good. Cryptography cannot make a model accurate. It only makes our reporting of that model honest, which is a different and much rarer property in this industry.
We would rather state these than have them found.
These are the changes we expect to make, published in advance so the record shows what changed and when:
Any change to the model produces a new model_version string, and every
prediction in the ledger records the version that produced it. A prediction can
always be traced to the exact model that made it.
This site publishes predictions for analysis and entertainment. We do not accept wagers, hold funds, or pay prizes. We are not affiliated with the NFL, any league, team, or sportsbook. Full disclaimers, including responsible-gambling resources, are at /disclaimers.
Last updated 2026-08-02. Model version elo-mov-v1+iso.