Evaluation & quality

Benchmark

A public, standardised test, with a fixed set of tasks and a fixed scoring method, used to compare AI models with one another.

scoretime →illustrative curvesmaximum scoreBenchmark A: saturatedBenchmark B: newer, harderTHE SAME TWO MODELSOn the leaderboardA is three places above BOn its own 150 documentsB 94% · A 88%the decision rests on thisA saturated benchmark no longer separates the leaders, and no leaderboard contains your documents.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/benchmark

In plain terms

A benchmark is a standardised exam for AI models. Every model sits the same questions under the same rules, so the scores can be placed side by side in a league table, usually called a leaderboard. Like school grades, the result says something about general ability and little about how the candidate will perform in one particular job.

Why it matters

Benchmark scores fill vendor announcements and procurement decks, so a manager needs to know how much weight they can carry. They are useful for cutting a long list of models down to a short one. They wear out in two ways. Saturation: once the best models all score near the maximum, the test no longer separates them. Contamination: public questions end up in training data, and a model scores well on items it has already seen. A leaderboard position therefore says little about a specific business task; a test on the organisation's own cases does.

Example

A logistics company must choose between two models for reading customs documents. Model A sits three places above Model B on a well-known leaderboard. On the company's own set of 150 documents, B extracts the fields correctly in 94% of cases and A in 88%, and B costs a third as much. The company chooses B.

Most often confused with

Benchmark vs. Evals

BenchmarkPublic and general; the same for every model and buyer
EvalsPrivate; written for one system and its tasks

The two differ in who sets the questions. Benchmark questions are written by researchers for everyone and stay fixed for years. Eval cases are written by one team for its own system and change with the product. A model can lead a benchmark and still fail an eval, because the benchmark never contained that company's documents, customers or rules.

Origin: The word comes from surveying, where a bench mark is a fixed point of known height, cut into stone, from which other heights are measured.

Under the hood

A benchmark consists of a dataset, a metric and a protocol for running models; results are published on leaderboards. Well-known examples: MMLU, multiple-choice questions across 57 subjects, on which leading models now score above 90%; SWE-bench Verified, 500 real software issues that a model must fix; ARC-AGI, abstract visual puzzles; Humanity's Last Exam, expert-level questions written to resist saturation; and Arena (formerly LMArena), which ranks models by anonymous head-to-head votes from users. Known problems: saturation; contamination, when test items leak into training data; overfitting to the test through repeated tuning, a case of Goodhart's law; differences in prompts, tools and number of attempts that make reported scores hard to compare; and errors in the items themselves. In February 2026 OpenAI stopped reporting SWE-bench Verified results, citing contamination and flawed test cases. Countermeasures include private held-out test sets, regularly refreshed questions and independent re-runs.

Written by Mehmet Erkek · Last updated: