The Model That Earns Its Keep
2026-08-16 · validation · analytics · projections
Is our projection actually better than a simple average? We ran the honest test - forecasting one player's WAR next season, walk-forward across thirteen seasons - and beat the baseline that is supposed to be unbeatable. Then we ran it on goalies, where we lose.
Figures computed 2026-09-13, against the models as they stood that day. We change them — see the methodology page for where they are now. Corrected 2026-09-13.
- Forecasting one skater's WAR: 0.624 against a three-year weighted average's 0.598 and last-season-only's 0.559.
- Walk-forward across 13 seasons and 6,976 player-seasons — each year projected by a model trained only on earlier years. We beat the average in 10 of 13.
- On goalies we lose. Our goalie projection ranks goalies worse than a three-year average does (0.024 vs 0.064), though its absolute error is lower.
- The skater margin is real and modest. Anyone claiming a chasm here is overfitting.
There is a quiet, uncomfortable question hanging over every fancy sports model: is it actually better than a simple average? It is uncomfortable because the answer is often no. If a sophisticated model cannot clear that bar, all the machinery is decoration.
The test
We ran a walk-forward backtest across thirteen seasons. For each season, the model is trained only on seasons that came before it, then projects every player, and we compare against what actually happened — against two baselines: last year only, and a 3/2/1 recency-weighted average of the last three seasons.
| Model | R² vs actual | Avg error |
|---|---|---|
| Our projection | 0.624 | 0.520 WAR |
| 3-year weighted average | 0.598 | 0.534 WAR |
| Last year only | 0.559 | 0.582 WAR |
Skaters with 40 or more games, 6,976 player-seasons from 2013-14 through 2025-26. Season by season the model beats the weighted average in ten of thirteen years, and the largest margin is the most recent one.
Why this is the test that matters
This is a player-level test, not a team-level one, and that is the point. When you sum a roster into one team number, individual accuracy washes out — the over- and under-estimates cancel, and everyone bumps into the same wall of unforeseeable trades, injuries and puck luck. One player at a time, there is nowhere to hide: if the model is just re-describing the average, it cannot beat the average.
The honest footnote
The margin is modest. 0.624 against 0.598 is a real and repeatable edge, not a chasm.
And on goalies, we lose:
| Model | R² vs actual | Avg error |
|---|---|---|
| Our goalie projection | 0.024 | 1.91 WAR |
| 3-year weighted average | 0.064 | 2.00 WAR |
| Last year only | 0.039 | 2.33 WAR |
Our goalie projection has less ability to rank goalies than a three-year average does. It has lower average error — it refuses to extrapolate a hot season and stays near the mean — but on the question of who will be better than whom, it is close to uninformative. Goalie value from one season to the next correlates at 0.209, so there is very little signal for anyone to find, which is why every public model lands near zero here. But "everyone else is also bad" is an explanation, not an excuse.
For skaters, the projection earns its keep. For goalies, our honest advice is to lean on the error bars and not the ordering.
Frequently Asked Questions
Is Hockey Alchemy’s player projection better than a simple average?
For skaters, yes. Forecasting one player’s WAR for the next season, the model scores R-squared 0.624 against a three-year weighted average’s 0.598 and last-season-only’s 0.559, with lower average error, across 6,976 player-seasons. It beats the weighted average in 10 of 13 seasons.
Can anyone project NHL goalies accurately?
No. Goalie value repeats from season to season at a correlation of about 0.209, so there is very little signal to find. Our own goalie projection scores R-squared 0.024, which is below a three-year weighted average’s 0.064 - we have lower absolute error but less ability to rank goalies.
What is a walk-forward backtest?
For each season, the model is trained only on seasons that came before it and then used to project that season. No figure is computed on data the model was fit on, which is what makes the comparison against a naive baseline meaningful.
More from the Notebook
Our 2026-27 NHL Projection, and Exactly How Much to Discount It
Carolina 111.6 points, Vancouver 79.1 - and luck alone gives the average team an 80% range 21.8 points wide, two-thirds of the league's entire spread. Our measured error says the honest range is closer to 32. The projection beats the honest baseline, and here is exactly how much of the table you should refuse to read as a ranking.
How Good Is Our WAR Model? We Ran Four Tests and Lost Two
We tested our WAR model against HockeyStats and Evolving Hockey on the four questions an honest player-value model has to answer, matched to each benchmark's published protocol. Summed to a team it loses both tests - it forecasts next season worse than simply reusing last season's standings. Measured per player it wins both, repeating at 0.77 against 0.62 and 0.46.
Talent vs Production: Why the Box Score Misleads
Two wingers score 20 goals; only one of them will do it again. The gap between what a player produced and the process underneath it is the most useful idea in hockey analytics - and it is the split our expected-goals and finishing models are built to make.