The whistle they didn't blow
The NBA grades its own officiating in the final two minutes of close games and publishes the results. Eighty-nine thousand graded calls later, the errors have a shape — and the question everyone wants answered turns out to be the one the data is worst at.
Most sports leagues respond to a bad call by saying nothing. The NBA reviews it, writes down whether its own official got it wrong, and publishes the answer with the official’s crew named on the front.
The Last Two Minute Report covers any game within three points at the two-minute mark. Every officiated event in that window gets a verdict: a correct call, a correct non-call, an incorrect call, or an incorrect non-call. Since 2014-15 that adds up to 89,052 graded decisions across 4,608 games — 6.9% of them marked wrong, about 1.33 per game.
That number is the one that gets quoted, and it’s the least interesting thing here. What the reports actually show is that officiating error has a shape.
Referees don’t blow the whistle wrong. They don’t blow it.
The four verdicts are wildly lopsided, and not in the direction the shouting suggests.
Every verdict the league has issued
Correct non-calls dominate, because most of what happens in a crowded final two minutes is contact the officials decide to let go.
Show the data
| Verdict | Calls |
|---|---|
| CNC | 61,113 |
| CC | 21,827 |
| INC | 5,282 |
| IC | 830 |
Source: NBA Last Two Minute Reports
There are 5,282 incorrect non-calls and only 830 incorrect calls — a ratio of about 6.4 to one. When an NBA official makes a mistake in the last two minutes, it is overwhelmingly a mistake of omission. They let something go that they shouldn’t have.
That runs against how the error is usually discussed. The complaint is nearly always about a whistle that was blown — the ticky-tack foul, the phantom call that decided a game. The league’s own review says those are rare. What is common is the foul nobody called.
There is a plausible reason, and it isn’t incompetence. In the last two minutes of a close game an official who blows a marginal whistle inserts themselves into the result, and one who swallows it does not. The incentive runs one way. The report counts the consequence.
Two rules stand apart
Break the errors out by what was being called and the distribution is not a gradient. It is a cliff.
How often each kind of call is graded wrong
Share of reviews marked incorrect, for call types with at least 500 reviews.
Show the data
| Call type | Incorrect (%) |
|---|---|
| Defense 3 Second | 43.3% |
| Turnover: Traveling | 37.4% |
| Away from Play | 9.3% |
| Loose Ball | 7.8% |
| Stoppage: Out-of-Bounds | 7.2% |
| Shooting | 6% |
| Offensive | 4.6% |
| Personal | 4.1% |
Labels shortened; most read “Foul: …” in the source.
Source: NBA Last Two Minute Reports
Foul: Defense 3 Second is graded wrong 43.3% of the time and Traveling 37.4%. Nothing else clears ten. The third-worst category, Away from Play, sits at 9.3%, and the fouls that make up the overwhelming bulk of the reviews are all between four and eight.
Two rules, then, that the league’s own reviewers say are called wrong something like two times in five — and every other rule called wrong about one time in twenty.
They have something in common. A shooting foul is a judgement about contact, and contact is where the officials are already looking. Defensive three seconds and traveling are neither judgement calls nor about contact: they are technical violations that require watching something else entirely — a defender’s feet in the paint, a ballhandler’s pivot foot — while three officials are watching the ball and the bodies around it.
The failure isn’t that these are hard to adjudicate. It is that nobody is looking.
The question everyone asks
Which referees are worst?
The reports name the crew for every game, so the question looks answerable. Rank the officials by the share of calls graded wrong in the games they worked, and you get a leaderboard with a two-and-a-half point spread between top and bottom.
That leaderboard is almost entirely a lie, and it is worth being precise about why.
An official in this sample works a few thousand graded calls. At a base rate near 6.79%, random variation alone moves an individual’s rate by about 0.4 points. Line up 46 officials and some will look good and some will look bad without any of them being different — that is what noise does to a ranked list.
So the test isn’t whether the spread exists. It’s whether the spread is bigger than chance would produce. Here it is: the observed spread across officials is 0.55 points where chance alone predicts 0.4. Real variation between officials exists. It is just very small — about 0.38 of a percentage point, on a base of 6.79%.
Officials at both ends, after adjusting for noise
Each rate pulled toward the league average by the share of its spread that is signal rather than luck. The four highest and four lowest of the 46 officials with 150+ reviewed games.
Show the data
| Official | Adjusted error rate (%) |
|---|---|
| Tom Washington | 7.39% |
| Tre Maddox | 7.3% |
| Eric Lewis | 7.21% |
| Derrick Collins | 7.17% |
| Gediminas Petraitis | 6.45% |
| John Goble | 6.44% |
| Mark Lindsay | 6.36% |
| Tyler Ford | 6.25% |
League average 6.79%.
Source: NBA Last Two Minute Reports
Adjusted, Tom Washington‘s 8.32% becomes 7.39% and Tyler Ford‘s 5.69% becomes 6.25%. The gap between the best and worst official in the league falls from 2.6 points to about 1.1 — which, at roughly nineteen graded calls a game, is a difference of about one call every five games.
What the data cannot do
Two limits, and neither is fixable by better arithmetic.
The first is that the reports never say which official made a call. Three officials work a game and the report grades the game, so every error is charged to all three. What looks like a referee’s error rate is a property of the games they were assigned, not of their own whistle — and assignments are not random, since the league sends its most trusted crews to its biggest games. The original open-source effort on this data tried to attribute calls to individuals by hand, which tells you how much the automated version is missing.
The second is selection. Only games within three points at the two-minute mark are reviewed. This is not a sample of NBA officiating; it is a sample of officiating under maximum pressure, in exactly the situations where a mistake is most consequential and most scrutinised. A blowout’s final two minutes never appear. Whatever the real error rate of an NBA game is, it isn’t 6.9%.
It has been getting better
Share of graded calls marked incorrect, by season
From 13% in 2014-15 to 5.1% in 2025-26.
Show the data
| Season | Incorrect (%) |
|---|---|
| 2015 | 13% |
| 2016 | 11.5% |
| 2017 | 8.6% |
| 2018 | 6.1% |
| 2019 | 6.6% |
| 2020 | 6.7% |
| 2021 | 6.1% |
| 2022 | 7.6% |
| 2023 | 6.1% |
| 2024 | 6.1% |
| 2025 | 4.5% |
| 2026 | 5.1% |
The seasons through 2017-18 were parsed from the original PDFs and later ones from the league's JSON feed, so small season-to-season steps aren't worth reading closely.
Source: NBA Last Two Minute Reports
The decline is large and most of it happens inside the early era rather than at the seam between the two data sources, which is the reason to believe it rather than blame it on the pipeline. Officials in the last two minutes of close games are graded wrong about half as often now as they were when the league started publishing.
Whether that is better officiating, or a review process that has drifted toward its own employees over a decade, is not something these documents can answer. They are produced by the organisation that employs the officials, reviewing in slow motion a decision someone made in real time with one look at it.
That is worth holding onto alongside the rest. The most remarkable thing about the Last Two Minute Report is not any number in it. It is that a league volunteered to keep score of its own mistakes in public, and then kept doing it for eleven years while the number went down.