The Model That Earns Its Keep

2026-08-16 · validation · analytics · projections

Is our projection actually better than a simple average? We ran the honest test - forecasting one player's WAR next season, walk-forward across thirteen seasons - and beat the baseline that is supposed to be unbeatable. Then we ran it on goalies, where we lose.

Figures computed 2026-09-13, against the models as they stood that day. We change them — see the methodology page for where they are now. Corrected 2026-09-13.

TL;DR

There is a quiet, uncomfortable question hanging over every fancy sports model: is it actually better than a simple average? It is uncomfortable because the answer is often no. If a sophisticated model cannot clear that bar, all the machinery is decoration.

The test

We ran a walk-forward backtest across thirteen seasons. For each season, the model is trained only on seasons that came before it, then projects every player, and we compare against what actually happened — against two baselines: last year only, and a 3/2/1 recency-weighted average of the last three seasons.

ModelR² vs actualAvg error
Our projection0.6240.520 WAR
3-year weighted average0.5980.534 WAR
Last year only0.5590.582 WAR

Skaters with 40 or more games, 6,976 player-seasons from 2013-14 through 2025-26. Season by season the model beats the weighted average in ten of thirteen years, and the largest margin is the most recent one.

Why this is the test that matters

This is a player-level test, not a team-level one, and that is the point. When you sum a roster into one team number, individual accuracy washes out — the over- and under-estimates cancel, and everyone bumps into the same wall of unforeseeable trades, injuries and puck luck. One player at a time, there is nowhere to hide: if the model is just re-describing the average, it cannot beat the average.

The honest footnote

The margin is modest. 0.624 against 0.598 is a real and repeatable edge, not a chasm.

And on goalies, we lose:

ModelR² vs actualAvg error
Our goalie projection0.0241.91 WAR
3-year weighted average0.0642.00 WAR
Last year only0.0392.33 WAR

Our goalie projection has less ability to rank goalies than a three-year average does. It has lower average error — it refuses to extrapolate a hot season and stays near the mean — but on the question of who will be better than whom, it is close to uninformative. Goalie value from one season to the next correlates at 0.209, so there is very little signal for anyone to find, which is why every public model lands near zero here. But "everyone else is also bad" is an explanation, not an excuse.

For skaters, the projection earns its keep. For goalies, our honest advice is to lean on the error bars and not the ordering.

Frequently Asked Questions

Is Hockey Alchemy’s player projection better than a simple average?

For skaters, yes. Forecasting one player’s WAR for the next season, the model scores R-squared 0.624 against a three-year weighted average’s 0.598 and last-season-only’s 0.559, with lower average error, across 6,976 player-seasons. It beats the weighted average in 10 of 13 seasons.

Can anyone project NHL goalies accurately?

No. Goalie value repeats from season to season at a correlation of about 0.209, so there is very little signal to find. Our own goalie projection scores R-squared 0.024, which is below a three-year weighted average’s 0.064 - we have lower absolute error but less ability to rank goalies.

What is a walk-forward backtest?

For each season, the model is trained only on seasons that came before it and then used to project that season. No figure is computed on data the model was fit on, which is what makes the comparison against a naive baseline meaningful.

More from the Notebook

Our 2026-27 NHL Projection, and Exactly How Much to Discount It

Carolina 111.6 points, Vancouver 79.1 - and luck alone gives the average team an 80% range 21.8 points wide, two-thirds of the league's entire spread. Our measured error says the honest range is closer to 32. The projection beats the honest baseline, and here is exactly how much of the table you should refuse to read as a ranking.

How Good Is Our WAR Model? We Ran Four Tests and Lost Two

We tested our WAR model against HockeyStats and Evolving Hockey on the four questions an honest player-value model has to answer, matched to each benchmark's published protocol. Summed to a team it loses both tests - it forecasts next season worse than simply reusing last season's standings. Measured per player it wins both, repeating at 0.77 against 0.62 and 0.46.

Talent vs Production: Why the Box Score Misleads

Two wingers score 20 goals; only one of them will do it again. The gap between what a player produced and the process underneath it is the most useful idea in hockey analytics - and it is the split our expected-goals and finishing models are built to make.