Methodology

Fair to the referees, and fair to the statistics.

Refereeing is one of the few jobs assessed live, by millions, in real time, on the worst possible evidence. RefCard tries to do the opposite: assess it slowly, on the official record, with the benefit of the doubt built in. It will not settle every argument — it is meant to start them from the truth.

Built on
KMI panel
the league's independent Key Match Incidents panel — not our opinion
Seasons graded
3
2023–24 onward
A grade of 0 means
league average
the same error rate as a typical official, for the matches he actually took
Accuracy %
never computed
we count confirmed errors, not a rate — there is no complete denominator

Where the data comes from

Every grade is built on the Premier League's independent Key Match Incidents (KMI) Panel — a five-person body of former players, former coaches, and one representative each from the league and the referees' body (Pro Ref). It reviews a selected set of decisions — the contested ones — and rules whether the on-field call and the VAR call were each correct.

This matters because it means we are not substituting our own opinion for the referee's. We don't count the decisions fans were angry about. We count only the ones an independent panel formally ruled were wrong. If an incident isn't on the panel's list, we treat it as correct — even if it was controversial.

Absence from the panel's list is not a clean sheet. Because the panel looks at a selected set, a decision it never reviewed is one we know nothing about — not one it approved. That cuts both ways and we hold to it in both directions: we never count an unreviewed decision as an error, and we never present "no decisions reviewed" as a record without a blemish.

How we see the panel's verdicts. The KMI Panel doesn't publish its rulings in a public data feed. They reach the public through Dale Johnson — now BBC Sport's refereeing and VAR correspondent, previously at ESPN — whose week-by-week Key Match Incidents reviews are our source: ESPN's panel write-ups for 2023–24 and 2024–25, and Dale Johnson's BBC Sport VAR review for 2025–26. Every incident we grade is traceable to one of those published reviews, including the on-field and VAR vote split where the panel gave one.

From 2026–27 the League publishes the docket itself. The Premier League now puts out the panel's complete round-by-round results — every reviewed decision with both votes, the lead official, and the panel's reasoning in its own words. That is a different kind of source from the three seasons before it, and the difference matters when you read a figure that spans them:

How a grade is built

Count the panel-confirmed errors

For each referee, in each season, we take every incident the KMI Panel ruled an error — missed interventions, incorrect interventions, below-threshold mistakes, and incorrect second yellows.

Weight by severity, not by paperwork

A clear-and-obvious error the panel says VAR should have caught counts in full (1.0); a below-threshold on-field mistake that didn't meet that bar — or an incorrect second yellow — counts half (0.5). Severity drives the weight, not whether a detailed vote was published for that incident. This is why the grade doesn't track a referee's raw error count one-for-one: two officials with the same number of errors can grade differently if one's were clear-and-obvious and the other's were marginal. So a profile can read "7 errors, expected ~3" (raw counts) while the grade, which halves the below-threshold ones, sits closer to average.

Compare to the flat league average

Errors become a rate per match — so a busy referee isn't punished for volume — and are measured against a single flat baseline: what a typical official did that same season. The league averaged 0.187 errors a match in 2023–24, 0.147 in 2024–25, and 0.166 in 2025–26. We use each season's own rate rather than one pooled number, because the sourcing and the panel's standards differ from year to year. Negative means fewer errors than that baseline; positive means more.

Shrink small samples — and know when not to rank

A referee with a handful of games can look brilliant or dreadful on luck alone, so every referee is pulled toward the league average by an amount set by how much real, repeatable difference the season actually shows. When a season's referees turn out to be genuinely indistinguishable once that sampling noise is removed — as in 2025–26 — we don't rank them at all: we show the raw counts and say so.

Counts pass, rates fail — when we'll name one referee

Everything on this site is either a count or a rate, and the two carry very different weight. The test is simple enough to apply without redoing the statistics:

A count is a census. Paul Tierney worked 101 matches in the VAR booth; there were 63 confirmed on-field errors last season; Brentford were on the wrong end of 7 of them. Nobody estimated those — we counted them. There is no error bar, so there is nothing to separate, and naming the top of a count is fair.

A rate is an estimate. Fouls per game, cards per foul, errors per match, minutes added per match — each is a sample average that would come out differently with a different set of matches. It carries an error band, and two referees can only be ranked against each other if the gap between them is bigger than that band.

Almost always, it isn't. Within a single season no referee is separable from the next on any style measure — the apparent leader changes depending on how many matches you require, which means the ordering is telling you about sample sizes rather than about refereeing. Across a full 26-season career the error shrinks a great deal, and still the top few sit inside each other's range.

So we split the claim in two. We will say "this referee differs from the norm" when the difference is bigger than its own error band — over a career that clears comfortably, and it is a real, repeatable trait: split a referee's matches in half and the two halves agree closely. We will not say "this referee does it most", because that requires separating him from the referee immediately behind him, and the data does not support it.

Where you see an ordered table on this site, the order is there so you can find a referee and see roughly where he sits — not to crown the top row. We say so on the table itself.

The final number is what's left: how a referee did relative to a typical official over the same number of matches. Negative is better than the league average; positive is worse; zero is exactly average.

Why there's no difficulty adjustment

It's tempting to grade referees against the difficulty of their fixtures — surely derbies and top-six clashes breed more mistakes. We tested that, hard, and it isn't true. Two independent checks agreed:

The finding is worth stating plainly: refereeing errors are not predictable from the character of a match. Big games and derbies are no more error-prone than a mid-table Tuesday night. So we don't pretend otherwise — every referee is measured against the same flat league average, and we'd rather say that than dress a flat number up as a difficulty model it isn't.

What the league trusts him with

Separately from the grade, each referee's profile shows the kind of fixtures Pro Ref assigns him — a marker of standing and trust, not a measure of accuracy. Unlike errors, this does separate referees sharply: some are handed the marquee games week after week, others rarely. We count, pooled across seasons, how many of a referee's matches are:

Both lists are fixed in advance, which is the point: they carry no hindsight from where teams happened to finish, so a club having a good season cannot retrospectively turn its fixtures into marquee ones.

The six clubs and the twelve derbies, in full

Big six: Arsenal, Chelsea, Liverpool, Manchester City, Manchester United, Tottenham.

Derbies: Arsenal–Tottenham, Liverpool–Everton, Manchester United–Manchester City, Newcastle–Sunderland, Crystal Palace–Brighton, Aston Villa–Wolves, Chelsea–Tottenham, Chelsea–Arsenal, West Ham–Tottenham, West Ham–Chelsea, West Ham–Arsenal, Liverpool–Manchester United.

A “marquee” match is any of the three. We report the share, the counts, and the single biggest fixture a referee took. It reads fairly at both ends: a light assignment load is a fact about the league's choices, not a mark against the official.

What we found — and won't hide

The honest headline is that Premier League referees are far more alike than the weekly outrage suggests. Most sit within a few thousandths of the average, well inside the margin of error. The real differences live at the very top and very bottom — and even those are smaller than you'd guess.

Where we're honest about the limits. A single season is a noisy read — year-to-year stability is only moderate, so a one-season swing is usually variance, not a real change in ability. That's why we show three-season trajectories, not just a snapshot.
On the 2023–24 data. That season's record gives confirmed errors without the per-incident votes, so its grade is flagged as built on a coarser source than the two that follow and is not compared across them.
On small samples. Referees with fewer than 15 matches in a season are shrunk hard toward the average and shown separately. Their grades are estimates with wide uncertainty, not verdicts.
On 2025–26. That is the season where the field turned out to be indistinguishable, so it is shown unranked — see above. The raw counts are real; we just don't turn a difference that small into a league table.

How we name an official — and when we won't

Naming a referee is easy: the match record says who refereed. Naming the VAR is not. Every VAR name we hold comes from a single secondary source, and one source is enough to credit an official but not to accuse one. So a name reaches an error surface only when a second, primary source agrees.

Every appointment is therefore in one of three states, and they are not interchangeable:

The principle underneath, which we apply in both directions: absence of confirmation is not evidence of error, but a primary source naming someone else is. Collapsing those two would either paralyse the record — treat every unchecked appointment as unusable — or launder a known-bad one into an accusation.

A contradicted error is never deleted. It stays in the published total and moves to an unattributed bucket. Dropping it would shrink the error record to make a sourcing problem disappear, which is the failure this rule exists to prevent. Of 72 VAR-layer errors, 54 sit on a verified appointment, 1 is unchecked, 1 is contradicted and attributed to no one, and 16 have no appointment on record at all.

One further limit, and it is the League's rather than ours: where an incident was an assistant referee's call, the published record names the role and never the person. We hold both assistants for every fixture and still do not name one, because nothing says which of the two raised the flag. Those incidents count for the match and the clubs, and are never charged to the referee.

Net xG Impact — what an error was worth, and which way it went

A confirmed error is a count. Net xG Impact says how much it was worth, in the one currency football already has: expected goals.

The sign, because three pages show one and it decides how you read them. A club's net xG is what errors gave it, minus what they took away:

Per incident the figure is signed, symmetric and sums to zero: whatever one club lost the other gained, so across the whole league the ledger nets out exactly. That is a property of the definition rather than a finding, and it is why a club with a large number of errors in its matches can still show a net near zero — a reason the site shows the counts for and against beside the net, never instead of it.

Two constants carry every xG figure on this site. Both are stated here because everything else is arithmetic on top of them:

A goal wrongly allowed or disallowed is not priced at all. Its value is the expected goals of the shot itself, and we hold match totals rather than shot-level data. A constant would invent a number and a zero would say the decision did not matter, so those errors are counted and carry no value.

An error we cannot value is held out of the xG, never entered as zero. A missed dismissal leaves no event in the match feed — a card that was never shown is not recorded — so there is no minute to price it from. Those errors keep their place in every count, and the surfaces that total xG say how many they set aside.

The figure prices the decision, not the consequence. It does not know whether the goal was scored, and it is not a claim that the result would have changed.

The style axes — how a referee runs a game

Style is descriptive. None of it measures accuracy: a referee who whistles often is not worse than one who lets play run, he is refereeing differently. Three axes reach a card, each pooled across a referee's whole career:

Why only three, and how the others were tested. Split a referee's own matches into two random halves at random and compare the two figures: an axis that describes the man scores near 1, an axis that mostly reflects which matches he happened to get scores near 0. That is split-half reliability, and it is the first of two tests every candidate axis had to pass. Measured across all 26 seasons:

AxisSplit-halfOn a card?
Whistle — fouls per game0.949yes
Tolerance — fouls per card0.911yes
Cards — yellows per game0.868yes
Home/away lean0.485no
Dissent share of cautions0.471no
Late-game share (after 70')0.269no
Added time per match0.130no
Time-wasting share0.116no

There is a clear gap between the three we print and everything else. "Tightens late" had been the lead fact on seven of ten cards in one round, on an axis that barely repeats between two halves of a referee's own matches — so it came off.

Two axes were dropped on the second test, not the first, and the distinction matters. Home/away lean and dissent share both repeat well enough (0.485 and 0.471, above several figures the site once printed). They failed a different question: whether a named referee can be told apart from the league at all. Of twenty referees with enough reasoned cautions to test, one sits above the league's dissent share by more than his own error bar and one below — eighteen are indistinguishable. A trait can be real and still be unprintable beside a man's name, and that is the same rule the rest of the site runs on: counts pass, rates fail.

Why never one season. Over a full career the two halves agree closely; over three seasons they barely agree at all. A three-season window measures the window, not the man, so every style figure runs across all 26 seasons we hold and the ordering on those pages is navigation.

How added time's "expected" is calculated

The added-time page compares each referee's real second-half stoppage time against what a transparent model predicts from the visible stoppages in his matches. Here is that model in full.

explains of added-time variance (R²)
Read these as averages, not measurements. Each number is how much added time tends to rise with one more of that event, across 658 matches. The VAR figure is the clearest example: we count VAR decisions, we don't time them — a 20-second check and a four-minute review both count as one. So "VAR ≈ +1.5 min" is an association, not the length of a check. Injuries and actual VAR-check durations aren't in the data, which is why the model leaves most of the variance unexplained.

Go deeper

For the league-wide breakdown of what kinds of mistakes referees actually make — missed by VAR, wrongly rejected at the monitor, below the clear-and-obvious threshold — and how the totals have moved across three seasons, see the Error Anatomy page.