Security & safety

Data Poisoning

An attack that plants manipulated examples in the data a model learns from or retrieves, so that the model later behaves as the attacker intends: making targeted errors, favouring an outcome or responding to a hidden trigger.

TRAINING DATA · illustrative2,000,000 labelled transactions400 poisoned records · 0.02%Monthly retrainingthe data is taken as it comesModellearns the pattern as “legitimate”Overall accuracy: unchangedTarget pattern: treated as “safe”The attacker never touches the system; the data is left where the training process will collect it.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/data-poisoning

In plain terms

A chef learns from a cookbook. Someone who cannot get into the kitchen slips a few altered recipes into the book. Months later the chef cooks them faithfully, with no idea that anything is wrong. Data poisoning attacks a model the same way, through what it studies. The attacker never touches the system; they place the material where the training process will collect it.

Why it matters

Models learn from data nobody has read in full: web crawls, public datasets, user feedback, downloaded models. Poison introduced there is hard to see, because the model performs normally on every standard test and misbehaves only in the cases the attacker chose. It is also expensive to remove once training is finished. For most companies the exposure sits in three places: the open-weight models and datasets they download, the data they fine-tune on, and the documents their RAG system retrieves. Checking provenance takes effort and narrows what can be used; that is the price.

Example

A bank retrains its fraud model every month on transactions that analysts have labelled. A fraud ring spends six months making 400 small transfers of one particular pattern and never disputes them, so they enter the data marked “legitimate”. Among 2 million examples they are 0.02%. After retraining, the model treats the pattern as safe and the ring moves large amounts through it. Overall accuracy on the bank's test set has not changed.

Most often confused with

Data Poisoning vs. Prompt Injection

Data PoisoningCorrupts what the model learns; the effect persists
Prompt InjectionHijacks what the model reads; the effect ends with the session

Both feed a model hostile text, at different moments. Prompt injection acts at run time and ends with the session; remove the content and the behaviour is gone. Data poisoning acts before use, during training or indexing, and its effect stays in the weights or the knowledge base for every user. The line blurs in RAG: a planted document is poisoning when it is stored and injection when it is retrieved.

Under the hood

Goals: degrade overall quality, cause specific errors on chosen inputs, or install a backdoor, behaviour that appears only when a trigger phrase or pattern is present. Entry points: public web pages that crawlers collect, open datasets and model repositories, third-party fine-tuning data, user feedback and ratings that feed retraining, documents indexed for retrieval, and an agent's long-term memory. The amount needed can be small: in a 2025 study by Anthropic, the UK AI Security Institute and the Alan Turing Institute, about 250 documents were enough to install a simple backdoor in language models of several sizes, however much clean data surrounded them. Defences: record the provenance of data and models, pin and verify dataset versions, filter outliers and duplicates, restrict who can write to training and retrieval stores, test for known trigger patterns, and compare behaviour before and after each retraining. Detection after the fact is unreliable, so prevention carries most of the weight. Tools for artists such as Nightshade use the same mechanism on purpose, against image models trained on their work without consent.

Written by Mehmet Erkek · Last updated: