In plain terms
A model is a snapshot of the world on the day its training data was collected. The world keeps moving: customers change their habits, new products appear, fraudsters change tactics, prices rise. The model keeps answering as if nothing had happened, and each month it is a little more wrong. Nothing breaks and no alarm sounds. Like the map of a growing city, it ages without anyone touching it.
Why it matters
It means that an AI system is never finished at launch. The accuracy measured in the pilot is a starting value with an expiry date, and the budget needs a line for monitoring, fresh test data and periodic updates. Hosted language models add a second form: the provider updates or retires the model version, and behaviour shifts although nothing changed on your side. Pinning a fixed version and re-running your evals before every switch keeps that under control, though only until the pinned version is retired.
Example
An online retailer's model routes support tickets to the right team. In January it routes 94% correctly. In March the company launches a subscription product, and tickets arrive of a kind the model has never seen. A monthly check on 300 freshly labelled tickets shows 91% in March and 89% in April, below the agreed threshold of 90%. The team retrains on recent tickets in May; in June the figure is back at 94%.
Most often confused with
Model Drift vs. Data drift
Data drift means that incoming data no longer looks like the training data: other customers, other wording, other products. It is one cause of model drift, and often it does no harm. Quality can also fall while the inputs look unchanged, because the right answer has changed (concept drift) or the provider has changed the model. Watching inputs gives early warning; only measuring outcomes shows whether the model still works.
Under the hood
Types: data or covariate drift (the input distribution shifts), label drift (the mix of outcomes shifts) and concept drift (the relationship between input and correct answer changes). Detection: track input statistics with measures such as the population stability index or a Kolmogorov-Smirnov test, track output distributions and, most reliably, measure accuracy on freshly labelled samples. Ground truth often arrives late, so proxy indicators are used in between. Responses: retraining on a schedule or on a trigger, refreshing the knowledge base, updating prompts. For language-model applications: call a dated model snapshot and avoid aliases that move to the newest version by themselves; follow the provider's deprecation notices; keep a golden dataset and re-run it on every model, prompt or retrieval change and at regular intervals; compare the result with the stored baseline. A 2023 study by Chen, Zaharia and Zou documented large behaviour changes between two versions of the same hosted models within three months.