In plain terms
A new employee does not get full signing authority on the first day. At first they prepare the work and someone checks all of it; later they handle routine cases alone and ask about the unusual ones; eventually they report at the end of the week. Autonomy levels put names on those stages for an AI system, so that the question “how much may it do alone?” gets a precise answer: for this task, level three.
Why it matters
The word “agent” says nothing about how much freedom a system has, and that freedom is what a board, an auditor or a regulator wants to know. A scale turns a vague debate into a decision per task: which level, with which limits, and what evidence would justify moving up. Higher levels save more labour and let errors travel further before anyone sees them. The right level depends on the task more than on the technology: how costly a mistake is, and whether it can be undone. Level numbers from different vendors are not comparable.
Example
An insurer sets the level for its claims agent task by task. Replies to customers stay at level 2: the agent drafts, a claims handler sends. Glass claims under €300 run at level 4: the agent pays, and a supervisor watches the daily figures and can stop it. Rejections and claims above €5,000 stay at level 3: a person approves each one. After six months of audits with an error rate under 1%, the €300 limit rises to €500.
Most often confused with
Autonomy Levels vs. Human-in-the-Loop (HITL)
Human-in-the-loop names a single arrangement: the system waits for a person's approval. Autonomy levels describe the whole range, from a system that only suggests to one that reports after the fact, and human-in-the-loop sits in the lower middle of it. Use the scale to choose a level for each task; use human-in-the-loop and human-on-the-loop to describe how oversight works at that level.
Origin: The idea of numbered levels is borrowed from SAE J3016, the standard for levels of driving automation that SAE International first published in 2014.
Under the hood
The idea of numbered levels comes from driving: SAE J3016 defines six levels of driving automation, from 0 to 5. Agents have no equivalent standard. Published proposals differ in what they grade: one from 2025 defines five levels by the user's role (operator, collaborator, consultant, approver, observer), while others grade by capability or by the length of unsupervised work. A scale is usable when each level states who initiates, who approves, what the system may touch and how it is stopped. In practice the level is set per type of action and enforced in the harness: permission modes, approval gates on tool calls with side effects, spending and step limits, allowlists and escalation rules. Inputs to the choice: the impact of an error, reversibility, volume, the error rate measured in evals, and how quickly a fault would be noticed. Autonomy is raised gradually, on evidence, and lowered after incidents. An older reference point is the ten-level automation scale of Sheridan and Verplank (1978).