Research What our model refuses to claim
What our model refuses to claim
Our NFL model beats its own predecessor on margin accuracy, loses to the closing spread in every held out season, and we print both numbers instead of picking the one that flatters us.
Here is the number most model marketing will not print. Across four held out NFL seasons, 2022 through 2025, the closing spread beats our own power rating on margin accuracy in all four, by about half a point of error. We built the model, we validated it, and the market still reads the game better than we do. That fact does not go in a footnote here. It is the lead, because a model that only tells you what it wins at is not being honest about what it is.
What the NFL model earns, and what it does not
Our current NFL rating, v3, replaced an older v1-implied baseline. On average game margin error over those four held out seasons (n = 1,087), v3 posts a mean absolute error of 10.000 points against the old baseline’s 10.425, a win in all four seasons, worth 0.425 points of MAE on average. That is real progress. The model is a better margin predictor than the one it replaced.
The closing spread’s error over the same games is 9.494, and it beats v3 in all four seasons. The spread is not a rough proxy here. Across the full 6,223-game set, the correlation between the closing line and the actual final margin is 0.433, which is why it is still the standard we grade against rather than a target we claim to have cleared.
Then there is the against the spread record, which is a different test entirely: 504-554 pooled over 1,058 graded games (pushes excluded), 47.64 percent, below a coin flip in three of the four held out seasons. Meanwhile the old v1 baseline, the model v3 beats on margin, actually grades better ATS: 50.76 percent over the same games. A model can be more accurate on average and worse against the spread at the same time, because ATS grading only asks which side of the number you landed on. A margin miss of half a point and a margin miss of nine points both lose the same way. So v3, if it fronts a public number, fronts expected wins and margin context. It does not front an against the spread edge, because it has not earned one.
The layer that mattered, and the ones that did not
The discipline that produces numbers like these is the same discipline that keeps features out of the model. We tested four additions to v3 on top of the base rating: a penalty for a team starting a quarterback other than its expected starter, a rest advantage term, a per team home field adjustment, and a roster age term. Rest advantage and per team home field failed outright and were dropped. Roster age cleared the rule by about 0.0013 points of margin error, small enough that we keep it in the formula but treat it as decoration and read nothing into its sign. Only the quarterback flag produced a real out of sample improvement. A non expected starting quarterback costs a team about 2.8 points, and adding that layer improved margin error on the 2022 to 2025 test set. A layer earns a permanent place in the model by clearing held out data, not by sounding plausible in a meeting.
The CFB composite: wins yes, market no
The college football preseason composite is a different kind of model, built to predict wins rather than single game margins, and it carries its own honest limit. Trained on 2021 through 2023 and tested out of sample on 2024 and 2025, it posts an average error of 1.880 wins per team against an SP plus carryover baseline at 1.966, and it wins the average without losing a season (2024: 1.844 versus 2.012, 2025: 1.915 versus 1.919, a tie inside a rounding error). That validates one thing: predicting actual wins better than the baseline it replaced.
It makes no claim about beating the market, for a specific reason. We do not have an archive of preseason win total lines stored far enough back to grade closing line value or return, so there is nothing to test that claim against yet. No market beating claim goes up until a full season of lines is stored and graded. The 1.9 win error bar is also why a small gap between our number and the market total is a lean, not a pick. A gap smaller than the model’s own average miss is noise. It takes a gap wider than 1.9 wins to clear that bar and become something we call a pick instead of a lean.
The number to respect
The closing spread is the sharpest number in football, built by people with more information and faster feedback than any model gets from a preseason build or a four season backtest. We grade against it. We do not claim to have beaten it, because on the numbers above, we have not. The record page carries every graded play against the closing line it was priced on, with nothing pulled after the fact. Members get the leans with the sample size and the error bar attached, not a headline stripped of both. Founding access is open at edgelabs.bet/join, and the full graded history sits at edgelabs.bet/record for anyone who wants to check the arithmetic themselves.