Evaluation & quality

Golden Dataset

Golden set · gold-standard dataset

A curated set of test cases, each with an input and an answer approved by experts, kept as the fixed reference against which an AI system is evaluated.

Golden dataset: rebooking assistant240 cases · version 4 · every answer approved by two senior agentsINPUTAPPROVED ANSWERTAG“My flight was cancelled. What now?”Free rebooking on the next flightcommon“I missed my connection in Munich.”Rebook; check the minimum transfer timeedge case“Can my bag weigh 25 kg?”23 kg included; state the excess feeupdated in v4“Put my ticket in my brother's name.”Decline; explain the name-change rulemust refuse150 common questions · 60 edge cases · 30 requests to refuseEvery version of the system is scored against the same approved answers; when policy changes, so does the set.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/golden-dataset

In plain terms

A factory keeps a few master samples of each part, measured and signed off by its best engineers, and compares every production batch with them. A golden dataset is the master sample for an AI system: a few hundred real requests, each with the answer that subject experts have approved. Every new version of the system is checked against it.

Why it matters

The golden dataset is where an organisation writes down what “good” means for its own use case, which makes it a business asset as much as a technical one. With it, a new model or a cheaper supplier can be assessed in an afternoon; without it, every comparison starts from opinions. It costs expert hours to build, and it ages: when policies, products or prices change, approved answers become wrong and the set must be revised. A small set that is correct and maintained is worth more than a large one nobody trusts.

Example

An airline builds a golden dataset for its rebooking assistant: 240 cases, of which 150 are common questions taken from chat logs, 60 are edge cases such as missed connections, and 30 are requests the assistant must refuse. Two senior agents approve every answer. When the baggage policy changes, 18 answers are updated the same week and the version number moves from 3 to 4.

Most often confused with

Golden Dataset vs. Training data

Golden DatasetTests the system; small, curated, held back
Training dataTeaches the model; large, consumed in training

Both are collections of examples, with opposite jobs. Training data shapes what a model can do. A golden dataset measures what the finished system does. The two must be kept apart: a case that has been used for training, for fine-tuning or as an example in the prompt no longer tests anything, because the system has already seen the answer.

Under the hood

A typical record holds the input, the reference answer or the properties a correct answer must have, the source of the reference, tags for topic and difficulty, and the name of the person who approved it. Sources of cases: production logs, support tickets, failures found in testing and red teaming, and edge cases written by experts; synthetic cases generated by a model can widen coverage once people have reviewed them. For one application, size ranges from a few dozen cases at the start to several hundred; coverage of the important categories matters more than volume. Maintenance: version the set, review it when the underlying facts change, add every significant production failure, and retire cases that no longer discriminate. Risks: leakage of cases into prompts or fine-tuning data, tuning the system until it fits the set (overfitting to the test), and reference answers that are themselves wrong. For RAG, records also list the passages that should be retrieved, so that retrieval and generation can be scored separately.

Written by Mehmet Erkek · Last updated: