Which Hockey Stats Are Skill, and Which Are Luck?

2026-08-16 · analytics · methodology · GAR · finishing

Line up every player's value in a category this season against next season and you get a brutal skill-vs-luck test. Staying out of the penalty box repeats more than four times better than finishing - and it changes how you should read a stat line.

TL;DR

Here is a simple, brutal test for any hockey stat: does it repeat? If a number measures real, durable skill, a player who is good at it this year should be good at it next year. If it is mostly luck, this year's value tells you almost nothing about next year's.

The skill-to-luck ladder

Year-over-year repeatability of each component, as a per-60-minute rate, across fifteen consecutive season-pairs:

ComponentRepeats atVerdict
Staying out of the penalty box0.57the stickiest thing we measure
Even-strength offense0.56skill, and by far the biggest component
Power-play offense0.42skill
Drawing penalties0.42skill
Even-strength defense0.42moderate
Penalty killing0.18weak, and tiny
Finishing0.12mostly luck

Read the top and the bottom against each other. A player's discipline is more than four times more repeatable than his finishing. Whether a skater takes penalties is close to a fixed trait. Finishing, the thing highlight reels obsess over, is nearly a coin flip from one season to the next.

Two notes before anyone over-reads that table

Penalty killing is the least important rung, not just the second-worst. Across all skaters in 2024-25 the spread in penalty-kill value is about a tenth of a goal, an order of magnitude smaller than any other component. A weak signal on a component that barely moves the total is a footnote.

Even-strength defense depends on which version you ask about. We estimate it three ways internally and they repeat at 0.20, 0.37 and 0.42. The number above is the one shown on the player pages. Defense is the hardest thing in hockey to attribute to an individual, and the fact that our three estimates disagree by that much is the honest measure of how hard.

Why this changes how you read a stat line

Be suspicious of any player evaluation that leans hard on last year's goal total. Goals are downstream of finishing, and finishing barely repeats. Meanwhile the player quietly generating chances and driving play at even strength is doing the thing that actually carries forward.

It also suggests the discipline column deserves more respect than it gets. It is not a big component, but it is the most reliable thing on the sheet: a player who took too many penalties last year is the safest bet on this list to do it again.

This is why our model is not points-based. It is built on the repeatable stuff — shot generation, on-ice impact, the tempo underneath the box score — and it deliberately discounts the parts that are mostly luck. When we build a player's total, we keep only half of his finishing, giving up a sliver of how well the model explains last season to buy a real gain in how well it measures the player.

Frequently Asked Questions

Which hockey stats are skill and which are luck?

Ranked by how well they repeat from one season to the next: staying out of the penalty box (0.57), even-strength offense (0.56), power-play offense (0.42), drawing penalties (0.42), even-strength defense (0.42), penalty killing (0.18) and finishing (0.12). The top of that list is close to a fixed trait; the bottom is close to a coin flip.

Does shooting percentage repeat year over year?

Barely. Finishing - goals scored above expected - is the least repeatable component we measure, at 0.12. Within a single season the signal is buried under variance, because a shooting percentage swings wildly on a few dozen shots.

Why does Hockey Alchemy only count half of a player’s finishing?

Because finishing barely repeats. Keeping all of it would make the metric better at describing last season and worse at measuring the player, so half is regressed toward what his chances say he should have scored.

More from The Lab

Talent vs Production: Why the Box Score Misleads

Two wingers score 20 goals; only one of them will do it again. The gap between what a player produced and the process underneath it is the most useful idea in hockey analytics - and it is the split our expected-goals and finishing models are built to make.

How Good Is Our WAR Model? We Ran Four Tests and Lost Two

We tested our WAR model against HockeyStats and Evolving Hockey on the four questions an honest player-value model has to answer, matched to each benchmark's published protocol. Summed to a team it loses both tests - it forecasts next season worse than simply reusing last season's standings. Measured per player it wins both, repeating at 0.77 against 0.62 and 0.46.

What's Actually Inside a GAR Number

GAR compresses a season into one figure - so the parts underneath had better add up. What our GAR is actually made of, why it is anchored on counting stats rather than built purely from RAPM (measured: pure RAPM drops standings R-squared from 0.75 to 0.51), and the bug that left our own component breakdown failing to sum to the number above it.