The Model That Earns Its Keep

2026-08-16 · analytics · methodology · projections · validation

Is our projection actually better than a simple average? We ran the honest test - forecasting one player's WAR next season, walk-forward across thirteen seasons - and beat the baseline that is supposed to be unbeatable. Then we ran it on goalies, where we lose.

TL;DR

There is a quiet, uncomfortable question hanging over every fancy sports model: is it actually better than a simple average? It is uncomfortable because the answer is often no. If a sophisticated model cannot clear that bar, all the machinery is decoration.

The test

We ran a walk-forward backtest across thirteen seasons. For each season, the model is trained only on seasons that came before it, then projects every player, and we compare against what actually happened — against two baselines: last year only, and a 3/2/1 recency-weighted average of the last three seasons.

ModelR² vs actualAvg error
Our projection0.6170.560 WAR
3-year weighted average0.5860.579 WAR
Last year only0.5540.624 WAR

Skaters with 40 or more games, 6,976 player-seasons from 2013-14 through 2025-26. Season by season the model beats the weighted average in ten of thirteen years, and the largest margin is the most recent one.

Why this is the test that matters

This is a player-level test, not a team-level one, and that is the point. When you sum a roster into one team number, individual accuracy washes out — the over- and under-estimates cancel, and everyone bumps into the same wall of unforeseeable trades, injuries and puck luck. One player at a time, there is nowhere to hide: if the model is just re-describing the average, it cannot beat the average.

The honest footnote

The margin is modest. 0.617 against 0.586 is a real and repeatable edge, not a chasm.

And on goalies, we lose:

ModelR² vs actualAvg error
Our goalie projection0.0211.82 WAR
3-year weighted average0.0631.94 WAR
Last year only0.0372.25 WAR

Our goalie projection has less ability to rank goalies than a three-year average does. It has lower average error — it refuses to extrapolate a hot season and stays near the mean — but on the question of who will be better than whom, it is close to uninformative. Goalie value from one season to the next correlates at 0.215, so there is very little signal for anyone to find, which is why every public model lands near zero here. But "everyone else is also bad" is an explanation, not an excuse.

For skaters, the projection earns its keep. For goalies, our honest advice is to lean on the error bars and not the ordering.

Frequently Asked Questions

Is Hockey Alchemy’s player projection better than a simple average?

For skaters, yes. Forecasting one player’s WAR for the next season, the model scores R-squared 0.617 against a three-year weighted average’s 0.586 and last-season-only’s 0.554, with lower average error, across 6,976 player-seasons. It beats the weighted average in 10 of 13 seasons.

Can anyone project NHL goalies accurately?

No. Goalie value repeats from season to season at a correlation of about 0.215, so there is very little signal to find. Our own goalie projection scores R-squared 0.021, which is below a three-year weighted average’s 0.063 - we have lower absolute error but less ability to rank goalies.

What is a walk-forward backtest?

For each season, the model is trained only on seasons that came before it and then used to project that season. No figure is computed on data the model was fit on, which is what makes the comparison against a naive baseline meaningful.

More from The Lab

How Good Is Our WAR Model? We Ran Four Tests and Lost Two

We tested our WAR model against HockeyStats and Evolving Hockey on the four questions an honest player-value model has to answer, matched to each benchmark's published protocol. Summed to a team it loses both tests - it forecasts next season worse than simply reusing last season's standings. Measured per player it wins both, repeating at 0.77 against 0.62 and 0.46.

Talent vs Production: Why the Box Score Misleads

Two wingers score 20 goals; only one of them will do it again. The gap between what a player produced and the process underneath it is the most useful idea in hockey analytics - and it is the split our expected-goals and finishing models are built to make.

Which Hockey Stats Are Skill, and Which Are Luck?

Line up every player's value in a category this season against next season and you get a brutal skill-vs-luck test. Staying out of the penalty box repeats more than four times better than finishing - and it changes how you should read a stat line.