Hockey + Technology / Field note
An Independent Hockey Analytics Audit Starts with a Decision, Not a Dashboard
A hockey analytics audit should answer a practical question: which decisions become more defensible with this
information, and where should staff withhold judgment? A polished dashboard can contain accurate calculations
while still answering the wrong question. Independent review starts with the intended action, then works backward
through the metric, its assumptions, and its source. Independence means testing those assumptions without a
preferred vendor or model to defend.
Write the decision contract first
Choose a recurring decision with an accountable owner. Specify when it occurs, what alternatives are available,
what evidence staff already use, and the cost of a wrong conclusion. An opponent-preparation meeting needs
different evidence from a technology procurement review. Define the output as something concrete: a reviewable
clip packet, a documented model limitation, or a recommendation to retain the existing process.
The contract should also state what the analysis cannot decide. A descriptive measure of territorial pressure
does not, by itself, establish which tactical adjustment caused that pressure or which adjustment will work next.
Agree on the distinction between describing events, predicting outcomes, and estimating the effect of an action
before choosing an evaluation method.
Trace the metric back to its ingredients
NHL EDGE provides public tracking-based insights. In its February 11, 2025 explanation, AWS describes Ice Tilt
as a view of game momentum and discusses Opportunity Analysis and Projected Goal Rate at shot release. These are
useful examples of different analytical questions, not interchangeable definitions of team performance. The
public descriptions neither validate every downstream use nor provide access to the underlying tracking feeds.
For each metric actually under review, build a provenance record: provider, permitted use, source event IDs,
extraction date, units, coordinate convention, filters, transformations, and definition version. Record missing
events and corrections rather than silently dropping them. Preserve an input snapshot or a reproducible reference
where licensing permits. A reviewer should be able to explain a changed result as a data correction, a code
change, or a changed definition.
Separate context from a convenient explanation
Inspect performance by strength state, score situation, period, opposition, and relevant deployment context.
Check whether coverage or tagging quality differs across those slices. A sample dominated by one situation may
not support a conclusion about another. Report how many usable observations remain after filtering, without
presenting a tiny subgroup as a reliable pattern.
Then review representative video with someone responsible for the tactical question. A failed exit might reflect
pressure, an unavailable support route, a change in progress, or a deliberate low-risk choice. Those possibilities
are hypotheses to investigate, not excuses to select after seeing the result. Observational association alone
does not identify a tactical cause, and unavailable actions cannot become recommendations.
Validate on information the model has not already seen
Keep development and evaluation separate. For a forward-looking use, train or tune on earlier games and evaluate
on later games. Keep related clips or events from the same game together where splitting them would leak context.
Audit whether any feature includes information recorded after the decision time. Freeze definitions, preprocessing,
and model selection before opening the final holdout.
Match the test to the claim. A probability model needs calibration checks as well as a proper scoring measure
such as Brier score or log loss. A retrieval tool needs checks that relevant, correctly aligned clips are found.
Compare each tool with a simple baseline and the current staff workflow. Report uncertainty and failures by
context; a single average can conceal an unusable situation. Do not repeatedly tune against the holdout and still
call it unseen evidence.
Illustrative scenario: reviewing a breakout question
Consider an invented preparation exercise in which staff want to compare two permitted breakout options against
a particular forecheck. No team, player, or measured result is represented here. The audit would define the trigger,
identify eligible possessions, align events with licensed video, and inspect whether both options were genuinely
available. It would retain contradictory clips and uncertain tags, then test whether the observation persists
in a later group of games.
The useful conclusion might be a narrower coaching question rather than a preferred option. If the data cannot
identify support positioning consistently, the next step is an annotation test, not a stronger recommendation.
That limitation is a valuable audit finding because it prevents precision in presentation from becoming false
certainty in a meeting.
Deliver a decision, a boundary, and an owner
The final package should contain the decision contract, metric dictionary, provenance record, reproducible test,
failure examples, and a prioritized action register. Each action needs an owner and an acceptance check. Classify
findings as ready for the stated use, restricted to specified contexts, or not supported. Identify what would
trigger a new review, including changed feeds, definitions, or tactical use. The audit earns its place by making
the next decision clearer, including when the responsible answer is to collect better evidence first.
Public sources and scope
These sources provide tracking and metric context; the audit method above is independent consulting analysis.
Discuss an independent hockey analytics audit focused on one decision,
clear evidence requirements, and practical acceptance criteria.