How Good Is Our WAR Model? We Ran Four Tests and Lost Two
2026-08-16 · analytics · methodology · GAR · validation
We tested our WAR model against HockeyStats and Evolving Hockey on the four questions an honest player-value model has to answer, matched to each benchmark's published protocol. Summed to a team it loses both tests - it forecasts next season worse than simply reusing last season's standings. Measured per player it wins both, repeating at 0.77 against 0.62 and 0.46.
- Both team-level tests: we lose. Summed to a team, our value explains the standings at R² 0.70 against HockeyStats' 0.82, and forecasts next season at 0.21 against their 0.35.
- Forecasting is worse than that. Our team totals predict next season less well than simply reusing last season's standings (0.21 vs 0.27).
- Both player-level tests: we win, clearly. Year-over-year repeatability 0.77 against HockeyStats' 0.62 and Evolving Hockey's 0.46, on an identical set of players.
- The pattern is the point: value summed to a team loses signal; value measured per player keeps it.
Every hockey analytics site will tell you its numbers are good. Almost none will show you the test that could prove them wrong. So we ran ours against the best published benchmarks in the sport on four questions any honest player-value model has to answer. We lost two of them.
The four tests
| Test | What it asks | Ours | Benchmark | Result |
|---|---|---|---|---|
| Descriptive | Do team WARs explain this season's standings? | 0.696 | 0.82 | loss |
| Predictive | Do they forecast next season's standings? | 0.211 | 0.35 | loss |
| Repeatability | Do good players stay good? | 0.599 | 0.382 | win |
| Projection | Can it forecast one specific player? | 0.617 | 0.586 | win |
All four are R², so they sit on the same scale. The first two are computed on HockeyStats' published protocol — pooled across team-seasons, against standings points per 82 — so the comparison is like-for-like.
First, an admission about how we got here
An earlier draft of this post claimed parity on the descriptive test. It compared our average of per-season R² against HockeyStats' pooled R² and called 0.76 against 0.82 a tie. Those are not the same statistic: ours averages sixteen separate within-season fits, theirs runs one regression across every team-season at once. Pooled numbers are almost always lower, because they also have to explain the differences between seasons.
Recomputed on their protocol, our 0.76 became 0.696, and the tie became a loss. We had been flattering ourselves with a mismatched comparison — precisely the thing this post was written to complain about.
"Standings R²" is not one number
That mistake is easy to make, because one question has at least six defensible answers and almost nobody states which they are quoting. On our data: average of per-season fits, 0.760; the same on our counting-component set, 0.825; pooled across team-seasons, 0.696; pooled on the counting set, 0.773; pooled inside HockeyStats' 2010-2021 window, 0.594; and 2024-25 alone, 0.888. Every one is honest. The spread is 0.29 — wider than the gap between us and any competitor.
Where we lose: anything summed to a team
Ours explains 70% of the variance in points per 82; HockeyStats' explains 82%. Restricted to the same seasons they used, ours drops to 59%. That is a real gap.
Forecasting is worse, and the honest way to show it is against the dumbest possible baseline. HockeyStats publishes that baseline themselves: prior-year standings predict next-year standings at R² 0.23. On our data the same baseline scores 0.269. Our team-aggregated value scores 0.211. If you want to guess how a team will do next season, last season's table is a better guide than summing our player values.
Two things cause this, and only one is a limitation of hockey. Team point totals absorb shootout results, one-goal-game variance and the bounces of an 82-game season — that caps everyone in the low 0.80s. The second is our own design: our even-strength defense term measures a player relative to the others on the ice with him. That framing is what makes it good at separating teammates, and it quietly subtracts what distinguishes one team from another. We have measured the trade directly — a more relative version of the model is more repeatable per player and worse at explaining standings at the same time.
Where we win: anything measured per player
Repeatability lines up every player's value this season against the same player's value next season. Run on the identical set of players present in all three models in both seasons, ours repeats at 0.774, against HockeyStats' 0.618 and Evolving Hockey's 0.464. Across our own fifteen season-pairs it holds at 0.767. Two pairs is the limit of the matched three-way comparison, so read that as two seasons of evidence, not sixteen.
Projecting a single player's WAR for next season, we score 0.617 against a three-year weighted average's 0.586 across 6,976 player-seasons, beating it in 10 of 13 seasons, out of sample.
So what is the number for?
Our WAR is a strong instrument for ranking and tracking individual players — who is good, who stays good, who will be good next year. It is a weak instrument for reconstructing or forecasting team quality, and you should not use it that way. The model we would have to build to win the team tests is a model that is worse at the thing we actually use it for.
Competitor figures for the first two tests are published by HockeyStats at hockeystats.com/methodology/war. Repeatability figures for other models are computed by us from their published season files, not quoted from them.
Frequently Asked Questions
How accurate is Hockey Alchemy's WAR model?
It depends entirely on whether you are measuring players or teams. Summed to a team it explains the standings at R-squared 0.70 against HockeyStats' 0.82, and forecasts next season at 0.21 against their 0.35. Measured per player it repeats year over year at 0.77 against HockeyStats' 0.62 and Evolving Hockey's 0.46, and projects an individual player at 0.617 against a three-year weighted average's 0.586.
Can WAR predict next season's NHL standings?
Not well, and ours does it worse than the simplest alternative. Team-aggregated WAR forecasts next season at R-squared 0.21, while just reusing last season's standings scores 0.27 on the same data. If you want to know how a team will do next year, the standings table is a better guide than summing player values.
Why does a player-value model lose at the team level?
Two reasons. Team point totals absorb shootout results and one-goal-game variance that no sum of player value can explain, which caps every model. And our even-strength defense term measures a player relative to the others on the ice with him - that framing separates teammates well, but it differences away part of what distinguishes one team from another.
More from The Lab
The Model That Earns Its Keep
Is our projection actually better than a simple average? We ran the honest test - forecasting one player's WAR next season, walk-forward across thirteen seasons - and beat the baseline that is supposed to be unbeatable. Then we ran it on goalies, where we lose.
Which Hockey Stats Are Skill, and Which Are Luck?
Line up every player's value in a category this season against next season and you get a brutal skill-vs-luck test. Staying out of the penalty box repeats more than four times better than finishing - and it changes how you should read a stat line.
What's Actually Inside a GAR Number
GAR compresses a season into one figure - so the parts underneath had better add up. What our GAR is actually made of, why it is anchored on counting stats rather than built purely from RAPM (measured: pure RAPM drops standings R-squared from 0.75 to 0.51), and the bug that left our own component breakdown failing to sum to the number above it.