Model customization

Synthetic Data

Data that is generated artificially, by a model, a simulation or a set of rules, to stand in for or add to data collected from the real world.

400 real casesrare, too fewLanguage modelwrites 20,000 variationsReviewa tenth is rejectedMixed datareal + syntheticCatch rate62% → 81% · illustrativeTest setreal calls onlySynthetic data carries the errors and biases of the model that made it; always test on real data.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/synthetic-data

In plain terms

Pilots train in simulators because nobody can wait for a real engine fire to practise on. Synthetic data is a simulator for AI: examples that never happened but look as though they could have. Invented customer complaints, made-up invoices, simulated sensor readings, fictional patient records. They are produced to order, in any quantity, including the rare and awkward cases that real life supplies too seldom.

Why it matters

It answers three recurring problems: too few examples of the cases that matter most, such as fraud or defects; real data that is too sensitive to share with developers or suppliers; and the need for test sets covering situations that have not occurred yet. The risks are as concrete. Synthetic data carries the blind spots and biases of whatever generated it, so a model trained on it can become confidently wrong about the real world. It is also not automatically anonymous: a generator trained on real records can reproduce parts of them. Keep the final test set real.

Example

A bank's fraud model misses a new phone scam because it has seen only 400 real cases. The team has a language model write 20,000 varied call transcripts in the pattern of those 400. Analysts review a sample of 500 and reject about a tenth as unrealistic, and the prompt is corrected. Trained on the mix, the model catches 81% of real cases of the scam, up from 62%. The test set contains only real calls.

Most often confused with

Synthetic Data vs. Anonymised data

Synthetic DataNewly generated records that describe no real person
Anonymised dataReal records with the identifying details removed or masked

Anonymised data is real: each row still corresponds to someone, with names and numbers stripped out, and combinations of the remaining details can sometimes identify them again. Synthetic data is created afresh from the patterns in the original. That lowers the risk without removing it, because a generator can memorise and repeat unusual records. Under data-protection law, both count as anonymous only if individuals cannot realistically be re-identified, and that has to be tested.

Under the hood

Three ways to produce it: rules and templates; simulation, as used for robotics, driving and sensor data; and generative models, including language models prompted to write examples, variations or whole conversations. With language models, quality depends on seeding with real examples, deliberately varying persona, topic and difficulty, and filtering the output, since unguided generations cluster around a few typical cases. Common uses: augmenting rare classes, building evaluation sets and red-team prompts, producing teacher outputs for distillation, and generating test data for software. Checks: fidelity (does it match real distributions?), utility (does a model trained on it work on real data?) and privacy (can real records be recovered?). Differential privacy gives formal guarantees at some cost in fidelity. A known long-run risk is model collapse: research published in Nature in 2024 showed that models trained repeatedly on the output of earlier models lose the rare parts of the original distribution. Keep a record of which data is synthetic.

Written by Mehmet Erkek · Last updated: