In plain terms
Put a few small stickers on a stop sign. Every driver still sees a stop sign; a vision model may read a speed limit. The stickers sit exactly where the model's judgement is most fragile. An adversarial attack searches for such weak spots and builds an input that looks ordinary to people and means something else to the machine. Nobody has broken into the model. It has been shown something it was never prepared for.
Why it matters
A model's test accuracy describes its behaviour on honest inputs. Wherever someone gains by fooling it, as in fraud detection, content moderation, identity checks, spam filtering or malware scanning, inputs will be shaped against it, and accuracy under attack can be far lower than the reported figure. Ask suppliers for results under adversarial testing as well as average accuracy. No model is fully robust, and hardening usually costs some accuracy on normal inputs, so the system around the model has to limit what one wrong decision can do.
Example
A marketplace uses an image model to block listings of counterfeit handbags; it catches 94% of them. Sellers of counterfeits discover that a faint pattern laid over the photo, invisible to shoppers, makes the model label the same bag “genuine”. Within two months the catch rate on new listings falls to 61%. The team retrains on manipulated photos, adds a check of seller history and sends borderline listings to reviewers.
Most often confused with
Adversarial Attack vs. Jailbreak
In the broad sense a jailbreak is one kind of adversarial attack, aimed at a language model's safety rules. In everyday use the term points at the older problem: classifiers and detectors fooled by engineered inputs, with no conversation involved. The difference matters for testing. A fraud model needs robustness tests on manipulated inputs; a chatbot needs red teaming of its conversations.
Origin: Adversarial examples were described in 2013 by Christian Szegedy and colleagues, who showed that imperceptible changes to an image could change a neural network's answer.
Under the hood
The standard taxonomy, used in NIST's report on adversarial machine learning, sorts attacks by stage and goal. Evasion: manipulate the input at the moment of use (adversarial examples). Poisoning: manipulate the training data. Privacy attacks: extract training data or infer whether a record was part of it. Model extraction: copy a model by querying it. An attack is white-box when the attacker knows the weights and can follow gradients, as in FGSM and PGD, and black-box when only the outputs are visible; examples built against one model often transfer to another. Physical versions exist: printed patches, altered signs, patterned glasses. For language models the equivalents are automatically generated adversarial suffixes and manipulated documents. Defences: adversarial training on attacked examples, input preprocessing, ensembles, detection of anomalous inputs, rate limits on queries, and human review where a decision matters. Robustness is measured as accuracy under a defined attack budget, and it is reported far less often than plain accuracy.