Evaluation & quality

LLM-as-a-Judge

LLM judge · model-graded evaluation

An evaluation method in which a language model grades the output of another AI system against written criteria, so that open-ended answers can be scored at a volume human reviewers cannot reach.

Rubricfour criteria, pass or failAnswer under testa claim summarySourcethe claim file it summarisesJudge modelapplies the rubricVerdictpass / fail + reasoncalibrates the judgeCalibration: 300 summaries already graded by expertsthe judge agrees with the experts on 91% of themA judge model is trusted only as far as it agrees with people on the same cases.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/llm-as-a-judge

In plain terms

An essay cannot be marked with an answer key; someone has to read it. When there are fifty thousand essays, a school hires assistant markers and gives them a marking scheme. LLM-as-a-judge does the same with a model: it receives the question, the answer and the scheme, and returns a verdict with a short justification. As with human assistants, the head teacher still spot-checks their marking.

Why it matters

Most valuable AI output is free text, where no exact-match check can say whether an answer is correct, complete or polite. A judge model makes it affordable to grade every test case on every change and to sample production traffic daily. The limit is that the judge is itself a model, with its own errors and biases. Until its verdicts have been compared with human judgements on the same cases, its scores are an opinion of unknown quality, and a dashboard built on them can mislead with great precision.

Example

An insurer's assistant writes 5,000 claim summaries a week. Claims experts can review about 100. The team writes a four-point rubric and has a judge model grade every summary. On 300 summaries that experts had already graded, the judge agrees with them in 91% of cases. Most of the disagreements concern missing policy numbers, so that criterion is rewritten and the comparison repeated.

Most often confused with

LLM-as-a-Judge vs. Human evaluation

LLM-as-a-JudgeA model applies the rubric: fast, cheap, at any volume
Human evaluationPeople apply the rubric: slow, costly, the reference

The two work together. Human evaluation defines what good means and produces the labels a judge is checked against. The judge then extends that judgement to volumes no team could read. Where the stakes are high or the criteria are subtle, as in medical advice or brand tone, people keep the final word and the judge serves as a first filter.

Origin: The term was popularised by a 2023 paper by Lianmin Zheng and colleagues, “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena”.

Under the hood

The judge prompt contains the criteria, the material to assess and, where one exists, a reference answer. Formats: a pass or fail verdict per criterion, a score on a scale, or a pairwise comparison of two answers. Design rules that improve reliability: one criterion per call, binary or low-resolution scales with each level described, reasoning requested before the verdict, examples of graded answers, and structured output. Documented biases: position bias in pairwise comparison (countered by swapping the order), a preference for longer answers and, in some studies, favouring the judge's own outputs. Calibration: have people label a sample, measure agreement with the judge using accuracy or Cohen's kappa, revise the rubric, and repeat whenever the judge model or its prompt changes. Choosing a judge from a different model family than the system under test reduces shared blind spots. Judges are weak where they lack the knowledge to verify facts; such checks need reference answers or retrieval.

Written by Mehmet Erkek · Last updated: