Who judges the AI judge? Why human calibration still matters 

LLM-as-a-judge has become a common way to scale AI evaluation. It uses a model grader to score model outputs or agent behavior against a rubric, compare responses, and run evaluations repeatedly as systems change. 

That solves a real scale problem. It also creates a new question: Who checks the judge? 

Automated graders make frequent evaluation practical. But before teams can rely on their scores, they need to know that the grader is measuring what they actually care about.

The judge is still a model

A model grader is itself a model, so its judgments also need to be validated. It is another model making a judgment about what should count as a strong or weak response. 

That means it can make its own mistakes. A grader can misunderstand the rubric, systematically prefer certain kinds of answers, miss errors that require domain expertise, or disagree with people on subjective or nuanced criteria. 

Research on LLM-as-a-judge has also documented issues such as favoring responses based on their order or length rather than their quality, as well as sensitivity to how evaluation prompts and criteria are constructed.

The point is not that model graders are unreliable by definition. It is that the measurement system itself has to be evaluated.

OpenAI's PaperBench offers a useful example. To scale grading, the team built an LLM judge to evaluate model performance. But they also built a separate benchmark, JudgeEval, specifically to measure how well that automated judge agreed with human expert labels.

This becomes especially important when there is no single right answer, when the task requires specialized expertise, when errors have meaningful consequences, or when the grader's score guides training or release decisions.  If the grader is measuring the wrong thing, teams can end up optimizing the system against the wrong standard. For example, if a grader consistently rewards confident-sounding answers that contain subtle errors, a team may see its evaluation scores improve even as the actual quality of the AI gets worse. 

So the central question becomes: How closely does the AI judge track the human judgment you actually trust?

Human judgment becomes the reference point

If a grader needs to be checked, the next question is: checked against what?

For judgments that cannot be verified with a simple rule or known answer, teams often need a human reference point. People review a set of outputs against defined criteria, creating examples to test whether the grader makes similar judgments. 

Hamel Husain and Shreya Shankar put the goal simply:

“What ultimately matters is how well your judge aligns with human judgments.”

In practice, this means having people judge a representative set of examples, then asking the model grader to score those same examples. Teams can compare the results to see where the grader agrees with the human judgments, where it disagrees, and whether those differences point to a problem with the grader, the rubric, or the task itself. 

In practice, that human reference point is often captured in a human-labeled reference set. When those labels are treated as the trusted standard for comparison, teams may also refer to it as a gold set. The exact terminology varies, but the purpose is the same: to give the grader a human standard to measure against. 

But creating a human reference is only part of the problem. The next question is who is qualified to make those judgments in the first place. 

The harder the judgment, the more the person matters

Not every human judgment carries the same weight. The person making the judgment needs to understand the work well enough to evaluate it.

For some tasks, that bar is relatively low. If the question is whether a response followed a required format or included a known fact, expertise may not be necessary. But other outputs are harder to judge. Evaluating clinical reasoning, a legal analysis, a cybersecurity response, or specialized engineering work may require someone who understands the domain well enough to recognize errors that could look correct or reasonable to someone without that expertise. 

That is why Hamel Husain recommends starting the LLM-as-a-judge process with a principal domain expert. The expert reviews real examples, decides what passes and fails, and explains the reasoning behind those judgments. The goal is to establish what quality actually means before trying to automate it. 

As Husain puts it:

“I don’t think you can eliminate humans completely, because the LLM still needs to be aligned to something, and that something is usually a human.”

The broader principle is visible in OpenAI’s GDPval. To evaluate models on real-world professional work, the researchers recruited experienced professionals from the occupations being tested. The experts who created the tasks averaged more than 14 years of experience, and model outputs were graded by professionals from those same occupations. 

The harder the work is to judge, the more important it becomes to have people with the right expertise setting the standard. 

Disagreement is information, too

For more subjective or consequential judgments, teams may also want input from multiple qualified people to see whether the standard holds across experts. Even qualified experts will not always agree. In many cases, that disagreement is not noise. It can reveal something important about the task itself. 

A rubric may leave room for interpretation. A response may be strong on one dimension and weak on another. Two experts may apply different but defensible standards. And for subjective or safety-sensitive questions, there may simply be no single answer that every qualified person would choose.

Recent research on human preference data has found exactly this. Michael JQ Zhang and colleagues studied where human reviewers disagree and found that disagreement often comes from meaningful differences such as an unclear or underspecified task, response style, and diverging preferences, not simply reviewer error.  

That matters when human judgments calibrate a model grader. If several experts score the same item differently, simply taking the majority judgement can hide information about where the evaluation standard itself is unclear.

The better question is not always, “Which reviewer was right?” Sometimes it is, “Why did qualified people disagree?”

That disagreement can help teams refine a rubric, identify ambiguous cases, or understand where reasonable experts are applying different standards to the same task. A 2026 EACL study similarly describes differences in human judgments as common and often informative because they can reflect task subjectivity and ambiguity. 

For AI teams, then, a strong human reference is not necessarily one where everyone agrees. It is one where teams can see who made each judgment, where agreement breaks down, and what those differences mean.

That is a useful signal to preserve rather than average away. 

Human calibration does not mean humans grade everything forever

The goal is to use human judgment where it adds the most value: establishing the standard, testing whether the grader matches it, and checking that alignment again when something meaningful changes.

Anthropic describes this as a common pattern in AI evaluation. Model-based graders can handle open-ended tasks and run repeatedly, but they should be calibrated against human graders for accuracy. Human reviewers provide the expert judgment used to establish and periodically check that standard.

That human reference is not permanent. Models change. Prompts change. Rubrics evolve. New edge cases appear. As the system changes, teams need to periodically check whether the grader is still aligned with the judgments they trust.

What you need is enough trusted human judgment to know whether the grader is still tracking the standard you care about. As the evaluation system improves, the amount of human effort can decrease, while periodic human checks remain part of the process.

That creates a practical division of labor: people establish and periodically check the standard; model graders apply it repeatedly at much greater scale.

Human calibration keeps the judge honest

Human calibration is ultimately about confidence in the measurement: knowing when the grader is aligned with the judgment you trust, when it is drifting, and when the standard itself needs another look.

As AI takes on more specialized and consequential work, the question becomes less about whether humans belong in the loop and more about whose judgment should define what good looks like.

‍


Related Articles