Security & safety

Red Teaming

Deliberately attacking your own AI system, the way an adversary would, to find its failures before someone else does.

1 · Attacktry it as a hostile user would2 · Findrecord the attack that worked3 · Fixprompt, guardrail, permissions4 · Retestdid the fix hold?LOOPrepeat on every releaseThe aim is to find where the system breaks before an attacker does.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/red-teaming

In plain terms

Banks hire people to try to break into their own vaults. Red teaming does the same for an AI system: a team plays the attacker, tries to make the model leak data, break its rules or take harmful actions, and writes down everything that works so it can be fixed.

Why it matters

Ordinary testing checks that a system does what it should. Red teaming checks what it can be made to do, which is where the expensive incidents come from. Regulators and enterprise customers increasingly expect evidence of it, and it is the only honest way to learn how your guardrails hold up under pressure.

Example

Before launching a customer-service agent, a team spends a week on it: planting instructions in test emails, asking for other customers' orders, coaxing it into refunds outside policy. They find eleven working attacks, fix nine and add human approval for the remaining two cases.

Most often confused with

Red Teaming vs. Evals

Red TeamingSearches for unknown failures
EvalsMeasures known behaviour

An eval measures performance against a fixed set of cases: does the system still get these right? Red teaming is open-ended and adversarial: what can we make it do that nobody listed? Findings from red teaming usually become new eval cases, so the same hole cannot reopen unnoticed.

Under the hood

Scope covers the model (jailbreaks, harmful content, bias), the application (prompt injection, data exfiltration, privilege escalation through tools) and the surrounding process. Methods combine manual expert probing with automated attack generation, in which one model produces and mutates attacks against another. Good practice: a written threat model, success criteria defined by impact, attack success rate tracked across releases, and retesting after every model, prompt or tool change. Reference frameworks include the OWASP Top 10 for LLM applications, MITRE ATLAS and the NIST AI Risk Management Framework.

Written by Mehmet Erkek · Last updated: