In plain terms
A wine producer can measure sugar and acidity in a laboratory and still employs tasters, because whether the wine is good is a human judgement. AI output is similar. Software can check that an answer has the right format; whether it is correct, helpful and in the right tone is decided by people who read it. Human evaluation organises that reading, with fixed criteria and several readers, so that the result can be counted and compared.
Why it matters
Human judgement is the reference for everything else in evaluation: automatic metrics and judge models are trusted only as far as they agree with people. It is indispensable when a product launches, when the stakes are high, when quality is a matter of expertise or taste, and when an automated score and user complaints point in different directions. It is also slow and costly, and people disagree with each other, tire and drift. The practical answer is to spend human attention where it counts most and to use it to calibrate cheaper methods.
Example
A bank prepares to launch an investment-information assistant. Three compliance officers each rate the same 200 answers as acceptable or unacceptable. They are unanimous on 122. After the rubric is rewritten with examples of borderline cases, a second round on 200 new answers brings unanimous agreement to 176. The exercise takes about 30 hours of expert time and yields the labels used to calibrate a judge model.
Most often confused with
Human Evaluation vs. Human-in-the-Loop (HITL)
Both put a person in front of AI output, for different purposes. Human evaluation is measurement: people score a sample, usually before launch or on a schedule, and the result guides development. Human-in-the-loop is control: a person approves a specific action in live operation before it happens. A system can pass human evaluation and still need a human in the loop for risky actions.
Under the hood
Three groups of evaluators: subject experts (few and expensive, required where correctness takes knowledge), trained annotators working from guidelines (larger volumes, consistent criteria), and end users (thumbs up or down, edits, A/B tests; plentiful and noisy). Formats: absolute rating against a rubric, pairwise preference between two outputs, ranking, and error annotation that marks what is wrong. Pairwise comparison is usually more reliable than a numeric scale. Good practice: written guidelines with examples, a calibration round, blind review in which raters cannot see which system produced an output, several raters per item, and random sampling. Inter-rater agreement is measured with Cohen's kappa for two raters and with Fleiss' kappa or Krippendorff's alpha for more; raw percentage agreement overstates consistency, because some agreement occurs by chance. Low agreement usually signals unclear criteria. Costs to plan for: recruitment and training, time per item, reviewer fatigue and exposure to disturbing content. The resulting labels serve to validate automatic metrics and judge models.