In plain terms
A satellite image suggests that a field is planted with wheat. To be sure, someone walks to the field and looks. What that person finds is the ground truth, and the image analysis is judged against it. In AI the idea is the same: for each test input there is an answer that has been checked by a trusted route, and the model's answer is right or wrong by comparison with it.
Why it matters
Every accuracy figure is a comparison with ground truth, so the figure can be no better than the reference behind it. The first question to ask of any reported score is who decided what was correct, and how. Reliable ground truth is often the most expensive part of an AI project, since it needs experts, measurements or waiting for real outcomes. For many tasks, such as summaries or advice, there is no single correct answer, and the reference becomes a set of criteria agreed by people, with their disagreements included.
Example
A hospital tests a model that flags fractures on X-rays, using 1,000 images. The ground truth for each image is the finding agreed by two radiologists reading independently. They differ on 64 images, which a third, senior radiologist decides. The model matches the ground truth on 930 images, an accuracy of 93%. The 64 disputed images are a reminder that the reference itself was built by people.
Most often confused with
Ground Truth vs. Golden dataset
Ground truth is the idea of a verified answer; a golden dataset is a set of test cases that each carry one. The first is the content, the second the container. The terms part ways outside evaluation: the labels in training data are also called ground truth, and so is the real outcome that a prediction is later checked against.
Origin: The term comes from remote sensing, where observations collected on the ground are used to check what aerial and satellite images appear to show.
Under the hood
In supervised learning, ground truth is the label attached to each training and test example; in evaluation, it is the reference each output is scored against. Sources, from strongest to weakest: direct measurement or system records, later real-world outcomes, consensus among several experts, a single annotator, and labels produced by a model (often called silver labels). Quality controls: written labelling guidelines, several annotators per item, inter-annotator agreement statistics, adjudication of disagreements, and audits of a sample. Known problems: label noise, which caps the score any model can reach and has been found in widely used benchmark datasets; ambiguity, where competent experts disagree; and staleness, when the world changes after labelling. Generative tasks often have many acceptable answers, so references take the form of required facts, rubrics or several example answers. Outcome-based ground truth can arrive late and only for some cases: a lender learns whether a loan was repaid only for the loans it granted.