In plain terms
A student who memorises last year's exam papers scores full marks on those papers and fails when the questions change. The student learned the answers and missed the subject. A model can do the same: with plenty of capacity and too little data it stores the training examples, quirks and errors included, and has little that carries over to tomorrow's cases. The warning sign is always the same: excellent results in the lab, disappointing ones in use.
Why it matters
It is the main reason a model that shone in the pilot disappoints in production, and the reason an accuracy figure from a supplier means little until you know which data it was measured on. The protection is discipline more than technology: keep a test set that neither the model nor its developers have used, and judge on that. The same applies when buying language models: a high score on a public benchmark whose questions leaked into the training data says little about your documents. The price is data set aside, never used for training.
Example
A bank trains a model to predict which customers will leave. After 40 rounds of training it is wrong about 1% of the 50,000 customers it learned from and about 26% of 5,000 customers it has never seen. The curves show why: error on the unseen customers was lowest at round 12, at 17%, and has risen since, while error on the training data kept falling. The team uses the version from round 12.
Most often confused with
Overfitting vs. Underfitting
They are the two ways of missing the right level of detail. An underfitted model is too simple or too briefly trained: it does badly on the training data and on new data. An overfitted model does well on the training data and badly on new data. The gap between the two scores is the test: poor scores with a small gap point to underfitting, a large gap to overfitting.
Under the hood
Diagnosis: compare training and validation error over time; a widening gap, with validation error turning upwards, is the signature. Causes: a model with too much capacity for the amount of data, too few or unrepresentative examples, training for too long, noisy labels, and data leakage between training and test sets. Remedies: more or more varied data, data augmentation, regularisation (L1 and L2 penalties, weight decay, dropout), early stopping, simpler models, ensembling and cross-validation. The split matters: a validation set for tuning and a separate test set used once, because repeated tuning against one test set overfits to it as well. In terms of the bias-variance trade-off, overfitting is high variance. Very large neural networks complicate the picture: they can fit the training data completely and still generalise (double descent). Related problems in language models: verbatim memorisation of training text, overfitting when fine-tuning on a few hundred examples, and benchmark contamination.