Security & safety

Data Exfiltration

The unauthorised transfer of data out of a system; with AI, typically an agent tricked into sending private information to an attacker.

ORGANISATION BOUNDARYPrivate dataemail, filesAgenttrickedEXIT CHANNELSsend emailHTTP requestimage URLpublic repoAttackergets the dataEach tool looks harmless; the danger is every channel through which data can leave.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/data-exfiltration

In plain terms

A thief needs two things: a way to reach the valuables and a way to carry them out. An AI assistant connected to your documents already has the first. Exfiltration is about the second: any channel through which the assistant can send something outside, whether an email, a web request or even a link.

Why it matters

Leaked customer data, source code or credentials are the costliest AI incidents, and the exit route is often a feature nobody thought of as risky. When reviewing an AI system, list every way information can leave it. That list, and not the model, is what determines exposure.

Example

A chat assistant with access to internal documents can display images. An injected instruction tells it to show an image from the attacker's server, with a summary of the confidential document packed into the image address. The picture never loads, but the request with the data has already been sent.

Most often confused with

Data Exfiltration vs. Data leakage

Data ExfiltrationDeliberate theft by an attacker
Data leakageAccidental exposure with no attacker

Leakage is unintentional: an employee pastes a contract into a public chatbot, or a model repeats something from its training data. Exfiltration is an attack: someone engineers the system into sending the data out. The controls overlap, but leakage is addressed mainly with policy and training, exfiltration with architecture.

Under the hood

Channels to audit: outbound HTTP from tools or code execution, rendered Markdown images and links, email and messaging tools, writes to shared or public locations (repositories, tickets, documents), DNS lookups, and URLs the user is asked to click. Controls: egress allowlists, disabling or proxying image rendering, blocking URLs that were not present in trusted input, human approval for outbound actions, data-loss-prevention scanning of outputs, and keeping secrets out of the context altogether. Exfiltration is the third leg of the lethal trifecta: remove the exit and injected instructions have nowhere to send what they find.

Written by Mehmet Erkek · Last updated: