In plain terms
Banks hire people to try to break into their own vaults. Red teaming does the same for an AI system: a team plays the attacker, tries to make the model leak data, break its rules or take harmful actions, and writes down everything that works so it can be fixed.
Why it matters
Ordinary testing checks that a system does what it should. Red teaming checks what it can be made to do, which is where the expensive incidents come from. Regulators and enterprise customers increasingly expect evidence of it, and it is the only honest way to learn how your guardrails hold up under pressure.
Example
Before launching a customer-service agent, a team spends a week on it: planting instructions in test emails, asking for other customers' orders, coaxing it into refunds outside policy. They find eleven working attacks, fix nine and add human approval for the remaining two cases.
Most often confused with
Red Teaming vs. Evals
An eval measures performance against a fixed set of cases: does the system still get these right? Red teaming is open-ended and adversarial: what can we make it do that nobody listed? Findings from red teaming usually become new eval cases, so the same hole cannot reopen unnoticed.
Under the hood
Scope covers the model (jailbreaks, harmful content, bias), the application (prompt injection, data exfiltration, privilege escalation through tools) and the surrounding process. Methods combine manual expert probing with automated attack generation, in which one model produces and mutates attacks against another. Good practice: a written threat model, success criteria defined by impact, attack success rate tracked across releases, and retesting after every model, prompt or tool change. Reference frameworks include the OWASP Top 10 for LLM applications, MITRE ATLAS and the NIST AI Risk Management Framework.