Security & safety

AI Safety

The field concerned with preventing AI systems from causing harm, whether through misuse by people, through failures of the system itself or through wider effects on society.

AI safetypreventing harm, whatever its causeMalicious usefraud, cyberattacksMalfunctionswrong output, loss of controlSystemic effectsjobs, dependence, concentrationThe working list for an organisation: risk assessment, evals, guardrails, human oversight, incident log.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/ai-safety

In plain terms

Aviation is safe because of everything around the engine: design rules, test flights, checklists, incident reports, inspectors. AI safety is the same idea applied to AI systems: the research, testing and rules meant to make sure that no serious harm follows when a system errs, when someone tries to abuse it, or when it turns out more capable than its makers expected. Alignment, training the model itself to behave well, is one part of that work.

Why it matters

Two levels concern a business. At the level of the model, safety is a selection criterion: developers differ in how they test before release, what they publish and how they respond to incidents. At the level of your own deployment, safety is your job: a safe model inside a careless system still causes harm. Regulators have moved the same way, and the EU AI Act places duties on providers and on deployers alike. Safety work costs time and sometimes capability, since safeguards also block some legitimate requests. The alternative is to find the failures in production.

Example

An insurer plans an assistant that answers customers' questions about health cover. Before launch a safety review asks three things: what happens when it is wrong, who could abuse it, and how failures will be noticed. A check of 300 test conversations finds 9 in which the assistant gives medical advice. The team narrows its scope, routes questions about symptoms to a nurse line, logs every refusal and names an owner for incidents. Launch moves by three weeks.

Most often confused with

AI Safety vs. AI security

AI SafetyKeeps the AI system from causing harm
AI securityKeeps attackers from damaging or hijacking the AI system

Safety is about harm that comes out of the system: wrong advice, dangerous content, actions nobody intended. Security is about attacks that go into it: prompt injection, data theft, stolen model weights. The two overlap, since an insecure system cannot be safe and a jailbreak belongs to both. They are usually run by different teams with different methods, and a project needs both reviews.

Under the hood

Risks are commonly grouped into three kinds, as in the International AI Safety Report: malicious use (fraud, cyberattacks, help with weapons), malfunctions (unreliable or biased output, loss of control over autonomous systems) and systemic effects (labour markets, concentration of power, dependence). Technical work includes alignment, robustness, interpretability, evaluations for dangerous capabilities and monitoring after release. Developers of frontier models publish safety frameworks that tie safeguards to capability thresholds, among them Anthropic's Responsible Scaling Policy, OpenAI's Preparedness Framework and Google DeepMind's Frontier Safety Framework, and report test results in model or system cards. Governments have set up institutes that test advanced models; the United Kingdom renamed its AI Safety Institute the AI Security Institute in 2025, a sign of how the two words mix in policy. For an organisation deploying AI the working set is smaller: a risk assessment for each use case, evals and red teaming before launch, guardrails, human oversight for consequential actions, incident logging and a named owner. NIST's AI Risk Management Framework and ISO/IEC 42001 give that work a structure.

Written by Mehmet Erkek · Last updated: