Security & safety

Guardrails

Checks placed around a model that inspect what goes in and what comes out, and block, rewrite or escalate anything that breaks the rules.

Inputuser, documentsInput guardinjection · PIIModelgeneratesOutput guardleaks · formatResponseor actionBlock · mask · route to a humananything a guard catches lands hereGuardrails lower risk without removing it: one layer among several, never the last line of defence.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/guardrails

In plain terms

A guardrail on a mountain road does not drive the car; it limits how wrong things can go. Around a model, guardrails are separate checks that sit before and after it: is this request in scope, does this answer contain personal data, is this action allowed? They work independently of whether the model behaves.

Why it matters

Guardrails are how a company's own rules, such as topics to avoid, data that must not leave and actions that need sign-off, are enforced on a model it did not train. They are necessary and not sufficient: each is a filter with an error rate, so they reduce incidents without guaranteeing safety, and they add latency and cost.

Example

A bank's assistant runs three checks on every turn: an input check that rejects investment-advice questions, an output check that masks account numbers, and an action check that sends any transfer above a limit to a human.

Most often confused with

Guardrails vs. Alignment

GuardrailsExternal checks around the model
AlignmentValues trained into the model

Alignment is about the model itself: training it to want the right things. Guardrails are outside the model: they catch problems whatever the model intends. A well-aligned model still needs guardrails for company-specific rules, and guardrails cannot make up for a model that works against them.

Under the hood

Types: input rails (injection and jailbreak detection, topic and PII filters), output rails (toxicity, leakage, groundedness and hallucination checks, schema validation), and action rails (allowlists, limits, approvals on tool calls). Implementations range from rules and regular expressions to small classifier models and LLM judges; deterministic checks belong in code, not in the prompt. Design questions: what happens on a block (refuse, redact, retry, escalate), how false positives are measured, and how much latency each layer adds. Instructions in the system prompt are guidance, not guardrails, because the model can be argued out of them.

Written by Mehmet Erkek · Last updated: