Methodology & Validation
← All methodologyAnyone can post a leaderboard of confident-looking numbers. Our difference is that every number ships with the uncertainty it actually carries, and nothing ships as a signal until it survives a gate designed to kill it. Here's exactly how that works — including the ideas it killed.
These aren't one-time checks. We re-run the full gate suite on every completed season — most recently the complete 2025-26 season. Metrics that were passing held; the ideas below that didn't clear the bar were re-tested and still don't clear it. A rejection here is a standing decision, not a stale footnote.
Every number ships with its uncertainty
A point estimate on its own invites false precision. So we resample whole games hundreds of times, recompute the whole rating each time, and report the 90% interval that produces. Read the point for the estimate; read the interval for how sure we are. Here is one rating shown the way we show all of them:
His impact is +0.263, and we are 90% confident it lands between +0.192 and +0.333. That whole interval clears zero, so he is distinguishable from a league-average driver — but a player whose interval overlapped his would be a statistical tie, and our leaderboards say so instead of inventing a rank.
See the "within noise" leaderboard →Does it hold up year over year?
A rating is only worth trusting if it repeats. Measured across six seasons, a skater's 5v5 offensive impact rank persists from one season to the next at a moderate, real level — around 0.45 — the signal you'd expect from a genuine talent measure rather than noise. Defensive impact is honestly noisier and persists less, around 0.30, and we say so plainly rather than dressing it up.
The intervals hold up too: for 93–95% of skaters, a season's 90% interval overlaps the next season's — so a year-to-year wiggle is usually within the noise the interval already told you about, not a real change. That's the honest version of consistency: the point estimates move a little, and the intervals were wide enough to have said they might.
These are out-of-sample persistence checks across the 2020-21 through 2025-26 seasons, not a curated single-season snapshot.
The gate every signal must clear
A candidate signal — a new feature, a new model, a clever-looking edge — doesn't get to ship because it looks good in a chart. It has to clear all five of these, in order:
- 1Beat the deployed model, not a strawman
A new idea has to out-predict the model we already ship, on data it has never seen — not a weakened baseline chosen to make it look good.
- 2A fresh held-out season
The test is a season the model was never trained on. In-sample gains are free and meaningless; only out-of-sample survival counts.
- 3Stable across 5 seeds
We refit across five random seeds. A result that only shows up on one lucky seed is noise wearing a result's clothes, and we drop it.
- 4A bootstrap interval that excludes zero
The lift gets a resampled confidence interval. If that interval touches zero, the effect isn't distinguishable from nothing — so it doesn't ship.
- 5A placebo noise-floor check
We run the same test on a scrambled, signal-free version of the feature. If the placebo scores as well as the real thing, the 'edge' was an artifact of the procedure, not the data.
What we tested and rejected — including our own ideas
The discipline only means something if it has teeth. It does. These are ideas that looked promising and did not survive the gate, so we don't ship them:
Some measures of the context a player skates in looked promising as inputs to our forecasts. Tested on top of the signals the model already uses, their added value was indistinguishable from zero and did not clear the gate. We kept them to describe a player's circumstances, not to predict one.
We tested whether our models produce a durable edge against the market. They don't — the price already subsumes what the model knows. So we don't sell picks or claim an edge. The product is honest analytics, not tips.
Goaltending is the noisiest thing to value in the sport. Our own testing puts game-level goalie performance at the noise floor for prediction, and single-season GSAx barely repeats year to year. We publish it as a leak-free scorecard of what happened, not as a signal for what's next.
We checked whether tagging a player 'declining' predicts a drop in next-season points. It doesn't — counting stats are sticky and luck props them up. So we don't present the label as a forecast, and we say so plainly on the metric itself.
What we don't claim
We don't sell betting picks or a pre-game "edge." We tested for one, and there isn't a reliable one — the market price already reflects what the model knows. What we're confident in, and will defend with the method above, is honest, uncertainty-quantified analytics: here's the number, here's exactly how sure we are, here's how it was built.
Validation figures reflect the data through 2025-26 and are recomputed each season. Head-to-head benchmark tests are dated snapshots of a specific held-out season, labeled where they appear on each metric's page.