Evaluation & quality

Ground Truth

Reference answer · gold label

The verified correct answer for a given input, established by measurement, records or expert judgement, against which the output of an AI system is compared.

WHERE GROUND TRUTH COMES FROMMeasurement and recordsa lab result, the system of recordExpert agreementtwo radiologists, a third decidesLater outcomewas the loan repaid?Model outputX-ray 417: “fracture”Ground truthX-ray 417: “no fracture”Comparethe same?No match: this image counts as an erroracross the whole test 930 of 1,000 images match: accuracy 93%A score can be no more reliable than the reference it is measured against.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/ground-truth

In plain terms

A satellite image suggests that a field is planted with wheat. To be sure, someone walks to the field and looks. What that person finds is the ground truth, and the image analysis is judged against it. In AI the idea is the same: for each test input there is an answer that has been checked by a trusted route, and the model's answer is right or wrong by comparison with it.

Why it matters

Every accuracy figure is a comparison with ground truth, so the figure can be no better than the reference behind it. The first question to ask of any reported score is who decided what was correct, and how. Reliable ground truth is often the most expensive part of an AI project, since it needs experts, measurements or waiting for real outcomes. For many tasks, such as summaries or advice, there is no single correct answer, and the reference becomes a set of criteria agreed by people, with their disagreements included.

Example

A hospital tests a model that flags fractures on X-rays, using 1,000 images. The ground truth for each image is the finding agreed by two radiologists reading independently. They differ on 64 images, which a third, senior radiologist decides. The model matches the ground truth on 930 images, an accuracy of 93%. The 64 disputed images are a reminder that the reference itself was built by people.

Most often confused with

Ground Truth vs. Golden dataset

Ground TruthThe correct answer for one case; a concept
Golden datasetA curated set of cases with their answers; a file

Ground truth is the idea of a verified answer; a golden dataset is a set of test cases that each carry one. The first is the content, the second the container. The terms part ways outside evaluation: the labels in training data are also called ground truth, and so is the real outcome that a prediction is later checked against.

Origin: The term comes from remote sensing, where observations collected on the ground are used to check what aerial and satellite images appear to show.

Under the hood

In supervised learning, ground truth is the label attached to each training and test example; in evaluation, it is the reference each output is scored against. Sources, from strongest to weakest: direct measurement or system records, later real-world outcomes, consensus among several experts, a single annotator, and labels produced by a model (often called silver labels). Quality controls: written labelling guidelines, several annotators per item, inter-annotator agreement statistics, adjudication of disagreements, and audits of a sample. Known problems: label noise, which caps the score any model can reach and has been found in widely used benchmark datasets; ambiguity, where competent experts disagree; and staleness, when the world changes after labelling. Generative tasks often have many acceptable answers, so references take the form of required facts, rubrics or several example answers. Outcome-based ground truth can arrive late and only for some cases: a lender learns whether a loan was repaid only for the loans it granted.

Written by Mehmet Erkek · Last updated: