In plain terms
An essay cannot be marked with an answer key; someone has to read it. When there are fifty thousand essays, a school hires assistant markers and gives them a marking scheme. LLM-as-a-judge does the same with a model: it receives the question, the answer and the scheme, and returns a verdict with a short justification. As with human assistants, the head teacher still spot-checks their marking.
Why it matters
Most valuable AI output is free text, where no exact-match check can say whether an answer is correct, complete or polite. A judge model makes it affordable to grade every test case on every change and to sample production traffic daily. The limit is that the judge is itself a model, with its own errors and biases. Until its verdicts have been compared with human judgements on the same cases, its scores are an opinion of unknown quality, and a dashboard built on them can mislead with great precision.
Example
An insurer's assistant writes 5,000 claim summaries a week. Claims experts can review about 100. The team writes a four-point rubric and has a judge model grade every summary. On 300 summaries that experts had already graded, the judge agrees with them in 91% of cases. Most of the disagreements concern missing policy numbers, so that criterion is rewritten and the comparison repeated.
Most often confused with
LLM-as-a-Judge vs. Human evaluation
The two work together. Human evaluation defines what good means and produces the labels a judge is checked against. The judge then extends that judgement to volumes no team could read. Where the stakes are high or the criteria are subtle, as in medical advice or brand tone, people keep the final word and the judge serves as a first filter.
Origin: The term was popularised by a 2023 paper by Lianmin Zheng and colleagues, “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena”.
Under the hood
The judge prompt contains the criteria, the material to assess and, where one exists, a reference answer. Formats: a pass or fail verdict per criterion, a score on a scale, or a pairwise comparison of two answers. Design rules that improve reliability: one criterion per call, binary or low-resolution scales with each level described, reasoning requested before the verdict, examples of graded answers, and structured output. Documented biases: position bias in pairwise comparison (countered by swapping the order), a preference for longer answers and, in some studies, favouring the judge's own outputs. Calibration: have people label a sample, measure agreement with the judge using accuracy or Cohen's kappa, revise the rubric, and repeat whenever the judge model or its prompt changes. Choosing a judge from a different model family than the system under test reduces shared blind spots. Judges are weak where they lack the knowledge to verify facts; such checks need reference answers or retrieval.