Security & safety

Adversarial Attack

Adversarial example

A deliberate attempt to make an AI model fail by giving it input crafted for that purpose, often a change too small for a person to notice that leads the model to a confident wrong answer.

Listing photomodel: “counterfeit bag”+A faint patterninvisible to the human eye=The same photomodel: “genuine”CATCH RATE FOR COUNTERFEIT LISTINGS · illustrativebefore94%two months on61%Nobody broke into the model: it was misled at the point where its judgement is most fragile.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/adversarial-attack

In plain terms

Put a few small stickers on a stop sign. Every driver still sees a stop sign; a vision model may read a speed limit. The stickers sit exactly where the model's judgement is most fragile. An adversarial attack searches for such weak spots and builds an input that looks ordinary to people and means something else to the machine. Nobody has broken into the model. It has been shown something it was never prepared for.

Why it matters

A model's test accuracy describes its behaviour on honest inputs. Wherever someone gains by fooling it, as in fraud detection, content moderation, identity checks, spam filtering or malware scanning, inputs will be shaped against it, and accuracy under attack can be far lower than the reported figure. Ask suppliers for results under adversarial testing as well as average accuracy. No model is fully robust, and hardening usually costs some accuracy on normal inputs, so the system around the model has to limit what one wrong decision can do.

Example

A marketplace uses an image model to block listings of counterfeit handbags; it catches 94% of them. Sellers of counterfeits discover that a faint pattern laid over the photo, invisible to shoppers, makes the model label the same bag “genuine”. Within two months the catch rate on new listings falls to 61%. The team retrains on manipulated photos, adds a check of seller history and sends borderline listings to reviewers.

Most often confused with

Adversarial Attack vs. Jailbreak

Adversarial AttackCrafted input makes the model misjudge what it sees
JailbreakCrafted conversation makes the model drop its rules

In the broad sense a jailbreak is one kind of adversarial attack, aimed at a language model's safety rules. In everyday use the term points at the older problem: classifiers and detectors fooled by engineered inputs, with no conversation involved. The difference matters for testing. A fraud model needs robustness tests on manipulated inputs; a chatbot needs red teaming of its conversations.

Origin: Adversarial examples were described in 2013 by Christian Szegedy and colleagues, who showed that imperceptible changes to an image could change a neural network's answer.

Under the hood

The standard taxonomy, used in NIST's report on adversarial machine learning, sorts attacks by stage and goal. Evasion: manipulate the input at the moment of use (adversarial examples). Poisoning: manipulate the training data. Privacy attacks: extract training data or infer whether a record was part of it. Model extraction: copy a model by querying it. An attack is white-box when the attacker knows the weights and can follow gradients, as in FGSM and PGD, and black-box when only the outputs are visible; examples built against one model often transfer to another. Physical versions exist: printed patches, altered signs, patterned glasses. For language models the equivalents are automatically generated adversarial suffixes and manipulated documents. Defences: adversarial training on attacked examples, input preprocessing, ensembles, detection of anomalous inputs, rate limits on queries, and human review where a decision matters. Robustness is measured as accuracy under a defined attack budget, and it is reported far less often than plain accuracy.

Written by Mehmet Erkek · Last updated: