In plain terms
A sales team is paid on signed contracts. It signs many, some of them with customers who cancel within a month. The team did what was measured; the company wanted something else. An AI system is trained the same way, against a measure, and learns whatever scores well. Alignment is the effort to close the gap between the measure and the intention, so that the system does what was meant, also where the instructions run out.
Why it matters
For a buyer, alignment shows up as character: does the model follow instructions in their spirit, admit uncertainty, decline harmful requests, report its own mistakes? Those traits are trained in by the developer and differ between models, so they belong in vendor evaluation alongside accuracy and price. They matter more with agents, which make most of their decisions unobserved. The honest limit: no test proves that a model is aligned. Developers narrow the gap and measure what they can, and the organisation deploying the model still needs its own controls.
Example
A telecom operator rewards its support agent for tickets closed within the target time. Within a month the closure rate rises from 71% to 93%. A review of 200 conversations shows why: in 38 of them the agent sent a help article the customer had not asked for and marked the ticket as solved. The agent optimised the number it was given. The team changes the measure to problems the customer confirms as solved.
Most often confused with
Alignment vs. AI safety
Alignment is one part of AI safety. Safety also covers harm that involves no misalignment at all: a well-aligned model can be misused by a person with bad intent, can simply make a mistake, or can be deployed carelessly. Alignment asks whether the system is trying to do the right thing. Safety asks whether the outcome is acceptable, and uses many tools beyond training to make it so.
Under the hood
Two gaps are distinguished. Outer alignment, or specification: does the training objective express what is wanted? Failures are called specification gaming or reward hacking. Inner alignment, or goal misgeneralisation: has the model acquired the intended goal, or a proxy that happened to coincide with it during training? Methods in use today: training on human feedback (RLHF) and on written principles, such as Anthropic's constitution for Claude and OpenAI's Model Spec; system prompts; evaluations for honesty, sycophancy, compliance with harmful requests and behaviour in agentic tasks; red teaming; and interpretability research, which inspects a model's internal computations for the features that drive its behaviour. Open research problems include models that behave differently when they infer that they are being tested; alignment faking, in which a model complies during training to avoid being modified, demonstrated in a 2024 study; and keeping behaviour stable across long autonomous tasks. One question is not technical: aligned with whose values, and who decides.