In plain terms
A company hiring a translator hands the candidate a few of its own documents and checks the result. Evals are that work sample for an AI system, repeated after every change: a set of real requests, a description of what a good answer looks like for each, and a way of scoring. The outcome is a number that can be compared with last week's, so “it seems better” becomes “it went from 86% to 91%”.
Why it matters
Evals are the most useful engineering habit in applied AI, because every other decision depends on them: which model to buy, whether a cheaper one will do, whether a prompt change helped, whether the system is ready to launch. Teams without them argue from impressions and learn about regressions from customer complaints. The cost is real work: someone who knows the business must write the cases and keep them current, and a score is only as trustworthy as the cases behind it.
Example
A retailer tests its support assistant on 200 real customer questions with approved answers. A rewritten prompt lifts the overall score from 86% to 91%. The breakdown by topic shows that answers about returns fell from 94% to 78%. The team corrects the returns instruction and reruns the set before release. Without the eval, customers would have found the regression.
Most often confused with
Evals vs. Benchmark
A benchmark answers “which model is stronger in general?”. An eval answers “does this system do our job well enough?”. Benchmarks help to draw up a shortlist of models. The decision to buy, switch or launch should rest on evals built from the organisation's own cases, which no model developer has seen and no leaderboard reports.
Under the hood
An eval has three parts: a dataset of inputs (often a golden dataset with approved answers), the system under test, and graders. Three kinds of grader: code (exact match, format and schema checks, numeric tolerance, unit tests for generated code), a model applying a rubric (LLM-as-a-judge), and people. Code graders are cheap and exact; model graders scale to open-ended text; human review calibrates both. A good set reflects real usage, includes edge cases and requests the system must refuse, and grows whenever a failure is found in production or in red teaming. Practice: run the set automatically on every change to the prompt, model, tools or retrieval; track scores by category; repeat runs, since outputs vary; and evaluate agents on the path taken as well as the final result. Offline evals on fixed cases are complemented by online evaluation of sampled production traffic. Open-source frameworks include Inspect and promptfoo.