Why Public Hockey Analytics Is Imperfect
2026-07-20 · analytics · commentary · methodology · expected goals
An honest accounting of where public hockey analytics hits a wall — the data ceiling (the feed records no passes, tracking drifts, private camera models see more) and the bigger problem now: a discourse that shares cards without meaning, rewards certainty over nuance, and stopped teaching. With the Zach Hyman case as the tell.
- Public hockey analytics is imperfect in two ways: the data has a hard ceiling, and the discourse around it has gotten lazy. The second problem is now the bigger one.
- The data limit is real and unfixable with public feeds: the play-by-play records no passes. Power-play xG is a wall for everyone (~0.72 AUC, ours and Evolving Hockey's alike), and camera-tracked private models see things public data never will.
- Hockey is the hardest of the major sports to model — few events, huge variance, thin public data. Our stats are genuinely predictive; they just explain less than baseball's, and that's the sport, not the analysts.
- Zach Hyman is the tell: nearly identical chance generation two years running (~49 xG), but 36 goals one year and 54 the next. The box score oversells him; the underlying numbers — and every honest fan — know it.
- The field has stopped teaching. Cards get shared as dunks, definitive one-liners beat honest uncertainty, and analysts have grown defensive instead of open. The fix is as much cultural as it is technical.
Public hockey analytics has quietly become the language fans, media, and even front offices use to argue about players. Expected goals, GAR and WAR, goals saved above expected — these numbers are on broadcasts and in contract debates, and for good reason: they are a real, measurable improvement over counting goals and eyeballing shifts. We build them, we publish them, and we stake our work on them.
This post is not a walk-back. It's an honest accounting of where public hockey analytics falls short — and it falls short in two very different ways. One is the data: the models are fenced in by a public feed that never recorded the things that matter most. The other is the discourse: the way analysts and fans talk about these numbers has gotten careless, absolutist, and strangely incurious. The data ceiling is the famous problem. The discourse is, at this point, the more damaging one — and the more fixable.
Part I — The data has a ceiling
Start with the limit everyone should already know. The NHL's public play-by-play logs events — shots, hits, faceoffs, giveaways, goals — with coordinates and timestamps. What it does not log is passes. There is no record of the cross-ice feed, the seam pass, the D-to-D setup that slides the goalie across the crease and turns a decent look into a near-certain goal.
That single absence is the hardest limit on expected goals. A one-timer off a cross-slot pass and a settled shot from the identical spot are wildly different chances, and to the public data they look nearly the same. We approximate the gap with prior-event features — rebound, rush, how far and fast the puck moved to get here — but an approximation of a pass is not a pass. Nowhere does it bite harder than the power play, where scoring is manufactured almost entirely through puck movement the data can't see. Our power-play discrimination sits around 0.72 AUC, and Evolving Hockey independently reports about the same. On pre-shot features alone, ~0.72 is the ceiling for all of us. A higher power-play number means a model has folded in information beyond the pre-shot chance itself — a legitimate design choice, just a different one than the line we draw. That's not a flaw in anyone's model; it's the edge of the data.
The passing gap is one face of a bigger one. The public feed sees the puck and little else — no continuous record of where the other nine skaters are or how the play was built. Camera-tracked private systems like Sportlogiq ingest full player and puck movement and grade the parts of the game the public feed never captured; that's a tier public analytics cannot fully reach. On top of that, the recorded data doesn't even sit still: shot locations and logging conventions drift season to season, so a model frozen on old data quietly loses calibration. Staying accurate is a retrain-against-drift discipline, not a one-time job — a model that hasn't been retrained lately isn't stable, it's stale. (Two smaller but real limits round it out: blocked shots are logged at the blocker's location, not the shot's, so every public model scores P(goal | unblocked shot) and drops the untrustworthy coordinates — see Distance and Angle Aren't Enough — and goalie GSAx, built on that same passless shot data, has near-zero team-level correlation, a limit shared across every public goalie model.)
How predictive is it, really?
Here's the part the loudest critics and the loudest boosters both get wrong. Public hockey analytics is predictive — meaningfully so. Our team GAR tracks team goal differential at an R² around 0.82 across eight seasons, and a team's expected goals predict its future goals better than its past goals do. That is the whole point of the exercise: strip the luck out, and what's left forecasts what happens next.
But be honest about the magnitude. Hockey is the hardest of the major sports to model, and it isn't close. Baseball is a sequence of discrete, isolated events — pitch by pitch, at-bat by at-bat — with enormous samples and Statcast measuring the speed and spin of every pitch and the launch angle of every batted ball; its advanced stats explain a lot and stabilize fast. Basketball has full optical tracking and a hundred possessions a night. Hockey is five or six goals a game of continuous, chaotic flow, with the thinnest public data of the three. Our models will always explain less season-to-season variance than baseball's — not because hockey analysts are worse, but because the sport is low-event, high-variance, and under-measured. A hockey model that claims baseball-grade certainty is lying about the sport it describes.
The Zach Hyman test
You don't need a model to feel this one — which is exactly why it's the perfect example. Zach Hyman scored 54 goals in 2023-24, and almost nobody, analyst or barstool, believes he's a 54-goal talent. Here's what our numbers say, and they tell the story cleanly:
| Season | GP | Goals | Individual xG | Finishing (G−xG) | GAR |
|---|---|---|---|---|---|
| 2021-22 | 76 | 27 | 30.0 | −3.0 | 10.7 |
| 2022-23 | 79 | 36 | 49.5 | −13.5 | 20.3 |
| 2023-24 | 80 | 54 | 49.3 | +4.7 | 26.6 |
| 2024-25 | 73 | 27 | 37.3 | −10.3 | 13.6 |
| 2025-26 | 58 | 31 | 35.4 | −4.4 | 15.0 |
Look at 2022-23 and 2023-24. Hyman generated nearly identical chances both years — about 49 individual expected goals each. But he scored 36 one year and 54 the next. That entire 18-goal swing is finishing: he under-converted by 13.5 in the first year and over-converted by 4.7 in the second, on the same volume of looks. The 54-goal season wasn't a different, better player — it was the same ~49-xG player having a hot finishing year on top of the most enviable setup in hockey.
This is analytics doing exactly what it should: it refuses to let the goal total tell the whole story, and it agrees with your eyes. But it also marks its own frontier. That ~49 xG is real — Hyman genuinely gets to the dangerous ice — but how much of it is Hyman versus Connor McDavid manufacturing the chances is a question the passless public data can only partly answer. Our isolated-impact models (RAPM, GAR) try, and they're honest that separating a net-front finisher from the superstar feeding him is one of the hardest problems public data poses. Hyman is a very good player. He is not a 54-goal one, and the numbers know it even as they admit what they can't fully untangle.
Part II — The discourse is the bigger problem now
The data ceiling is fixed and understood. What's gotten worse — and what actually holds the public game back — is how we talk about these numbers.
Certainty travels; nuance dies. The medium is partly to blame. In a character-limited post, "Player X is elite" fits and "Player X is a good-not-great driver with high on-ice variance and very favorable usage" does not. Definitive statements are useful when space is short — they're shareable, screenshot-able, argument-ending. So the format quietly rewards the confident one-liner and punishes the honest hedge, and analysts, wanting reach, drift toward absolutism the data doesn't support. The number was probabilistic; the tweet is a verdict.
Cards get shared without meaning. The player card — ours included, and JFresh's, and HockeyViz's — was supposed to be a conversation starter: here's a player's profile, now let's talk about what it does and doesn't capture. Instead it's become ammunition. A card gets screenshotted to win an argument, quote-tweeted as a dunk, waved around by people who couldn't tell you what a single bar on it means. Sharing went up; understanding didn't. A visualization that travels without its meaning isn't analysis — it's a flag people wave.
The field stopped teaching. If you publish a number, you owe the reader an account of what it means and what it doesn't — that's the deal. Somewhere along the way public analytics largely stopped honoring it. There's plenty of publishing and precious little explaining; the education that should travel with the metric got left behind. That responsibility falls on the analysts, and collectively we've let it slip. A public that misuses a stat is often a public that was never taught how to use it.
And the space got defensive. Healthy fields argue in good faith — they treat pushback as pressure-testing, not heresy. Too much of public hockey analytics now meets a question with a sneer, defends models instead of interrogating them, and mistakes confidence for rigor. That's the opposite of how the work should go. The best thing you can do for a model is attack it; the best thing an analyst can do is invite the attack.
Underneath all of it is a quieter problem: the public space has largely stopped progressing. The core public models are years old, the methods have plateaued, and genuinely new public work is rare — even as baseball and basketball raced ahead on richer data and open collaboration. Some of that is the data ceiling. A lot of it is a culture that started performing certainty instead of building the next thing.
The next frontier
Where does it go from here? On the data side, the frontier is obvious: public tracking. The NHL's Edge data is a real start — puck and player positioning, shot and skating speeds — but it's thin, partial, and largely locked away from the granular, shot-by-shot form a modeler needs. The day pre-shot passing and off-puck movement enter the public record is the day the power-play wall finally moves and the Hyman-versus-McDavid question gets a real answer. Until then, the honest public modeler works the edges: better prior-event features, cleaner isolation of individual impact, smarter handling of drift.
But the more urgent frontier is cultural, and it doesn't need better data at all. It needs analysts to teach again — to ship the explanation with the number. It needs less absolutism and more stated uncertainty. It needs the reflex to invite scrutiny rather than repel it, and to treat a card as the start of a conversation rather than the end of one. None of that is gated by the NHL's data feed. It's a choice.
What we owe
Public hockey analytics is a powerful lens with a fixed focal length, set by a feed that records no passes and drifts underfoot. On the data side, what we owe is straightforward: state exactly where the ceiling is, retrain relentlessly against drift, drop the shots whose coordinates can't be trusted, and validate honestly — out-of-sample, on matched tests when comparing models, never on a single number pulled from a friendly slice.
But the other half of the debt is bigger. We owe the reader the meaning, not just the metric. We owe uncertainty where it exists, an open door to the people poking holes, and the humility to say "here's what I can measure, and here's what I can't" — the way the Hyman numbers do. A model that tells you where it's weak is one you can actually reason with. A field that teaches what its numbers mean is one that earns the trust its confident one-liners keep borrowing. That's the work. The data will improve on the league's schedule; the culture can improve on ours.
For the full build — the four situation-specific models, the features, the head-to-head testing — see How Our Expected Goals Model Works and the methodology page.
Frequently Asked Questions
Is public hockey analytics reliable?
Yes, meaningfully — team GAR tracks goal differential at an R-squared around 0.82 and expected goals predict future goals better than past goals do. But it has a hard ceiling set by public data that records no passes and drifts season to season, and hockey is the hardest major sport to model. Treat the numbers as a strong lens, not settled truth.
Why is power-play expected goals so hard to model?
Because the NHL's public play-by-play records no passes. Power-play scoring is manufactured almost entirely by cross-ice puck movement the data never captures, so power-play xG is a wall for every public model — around 0.72 AUC for Hockey Alchemy and, independently, Evolving Hockey.
What does Zach Hyman's goal total show about box-score stats?
Hyman generated nearly identical chances in 2022-23 and 2023-24 (~49 individual expected goals each) but scored 36 goals then 54. The 18-goal swing was finishing variance on the same McDavid-fed chances — the 54-goal season was a hot-finishing version of the same player, not a better one.
More from The Lab
How Our Expected Goals (xG) Model Works
A transparent look inside Hockey Alchemy's xG model: situation-specific XGBoost models, 53 features, 16 seasons and 1.6M shots — and how it stacks up against MoneyPuck on a matched-shot test.
Distance and Angle Aren't Enough: Why Our xG Model Splits the Ice Into Zones
A pure distance-and-angle expected-goals model looks reasonable — and quietly misprices the whole ice. We map exactly where it goes wrong, then show how zone, rush, and prior-event features lift out-of-sample AUC from 0.70 to 0.84.
How We Calculate NHL Power Rankings: The Elo System Behind Hockey Alchemy
A transparent look at the Elo rating system that powers our NHL power rankings, game predictions, and Stanley Cup odds. K=12, home ice = 50 points, 30% regression to the mean — and three different ways we account for margin of victory.