Capability and autonomy are different questions
Most agent roadmaps chase one question: can the model do the task? A second question gets less attention. Should it do the task alone?
A model that can draft a customer reply can also send it. A model that can match a bank line to an invoice can also post the entry. Capability tells you what is possible. Autonomy tells you how much of it the product lets the agent do before a person looks.
Kevin Feng, David McDonald and Amy Zhang make this split explicit in their 2025 paper, Levels of Autonomy for AI Agents. They argue that an agent's level of autonomy is a deliberate design decision, separate from its capability. That framing hands the decision to product and design teams, which is where it belongs.
Picture autonomy as a dial rather than a switch. This article covers the five settings on that dial and how to choose one for each task. It then turns to the controls each setting needs.
The five user roles
Feng and colleagues define each level by the role the user plays. The labels below follow the paper, and the product examples are the ones the authors use.
- L1, Operator. The user plans and invokes every step. The agent supports on demand and acts when asked. Example: Microsoft Copilot.
- L2, Collaborator. User and agent plan together and split the work. The user can edit or take over the agent's work at any point. Example: OpenAI Operator.
- L3, Consultant. The agent plans and executes, and asks the user for expertise and preferences. The user steers by answering questions and requesting changes. Example: Gemini Deep Research.
- L4, Approver. The agent works alone and asks for help at blockers or before consequential actions. The user approves or rejects, and can set approval conditions in advance. Example: Devin.
- L5, Observer. The agent plans, acts and resolves problems alone. The user watches activity logs and holds an emergency off switch. Example: Voyager.
Two things stand out. The user's means of steering shrinks as the level rises, from full control at L1 to an off switch at L5. And the interface work changes shape at each level. At L1 you design invocation. At L4 you design an approval moment that has to carry the weight of everything the agent did before it.
The Cloud Security Alliance's Agentic AI Autonomy Levels and Control Framework (March 2026) uses a six-level scale for security teams and states that full autonomy is not recommended for current deployments. If your security or compliance team uses the CSA scale, map the two before you start. The user-role framing is easier to test with users. The CSA scale is easier to audit.
Set the dial per task, not per product
A common mistake is to give a whole product one autonomy level. Real workflows mix low-risk and high-risk steps, and each deserves its own setting.
Take a small business that handles customer service on WhatsApp:
- Answering opening hours or delivery areas. Low risk, easy to correct. The agent can reply on its own (L4 or L5), with a log the team reviews.
- Quoting a price or a delivery date. A wrong answer creates a commitment. The agent drafts and a staff member approves (L4), or the agent drafts inside the staff member's inbox (L2).
- Handling a refund or a complaint. Money and reputation are at stake. The staff member leads, and the agent summarises the history and suggests a reply (L1).
A finance team reconciling accounts looks similar. Matching bank lines to invoices can run at L3, with the agent asking about unclear items. Posting journal entries sits at L4 with an approval. Releasing a payment stays with a person.
Two questions that set the level
For each task, ask:
- What does a wrong action cost, and can someone undo it? A draft email costs nothing until it is sent. A sent email, a posted journal or a transfer cannot be recalled with one tap.
- How often is the agent right on this task, measured on real cases? Measure it with evals on your own data, not a demo or a vendor benchmark. Hamel Husain and Shreya Shankar's evals FAQ is a good place to start: review real traces, label the failures, then build pass or fail checks.
Low cost and high measured accuracy point to a higher level. Irreversible actions or thin evidence point to a lower one. When the answers conflict, the cost of the wrong action wins.
Design the controls for each level
Operator and collaborator (L1 and L2)
The design problem here is friction. People should be able to call the agent up without effort and dismiss it without penalty. Microsoft's Guidelines for Human-AI Interaction list both as separate guidelines: support efficient invocation and support efficient dismissal.
At L2, handover is the moment to design. When the agent passes work back, show what it finished and flag anything it was unsure about. When the person takes over mid-task, keep their edits. Nothing erodes collaboration faster than an agent that overwrites a correction.
Consultant (L3)
At L3 the agent leads and interrupts the user with questions. Each question costs attention, so each one should earn its place. Good consultant-level questions come early, bundle related choices together and explain why the answer matters. A research agent that shows its plan and asks "Include competitor pricing from Europe as well as the US?" before it starts saves a wasted hour. The same question at minute forty wastes one.
Approver (L4)
Approval is where most agent products fail their users. An approval screen that shows "Agent wants to send 14 emails. Approve?" invites a rubber stamp. Osmani calls this pattern cognitive surrender: once the loop runs itself, people stop forming their own opinion and accept what comes back.
Design approvals that make judgment possible:
- Show the consequence. Preview the exact email, the ledger entry or the calendar change. The HAX guidelines ask products to convey the consequences of user actions, and an approval is a user action.
- Show what is unusual. Highlight the items that differ from the pattern: the invoice with no matching purchase order, the reply that quotes a price.
- Keep batches small or split them. Let people approve routine items as a group and review exceptions one by one.
- Let people set conditions in advance. Feng and colleagues describe users pre-specifying when the agent must ask. "Always ask before any payment above USD 5,000" is a better control than a daily pile of approvals.
Birgitta Böckeler's harness engineering essay puts the aim well: direct human input "to where our input is most important." An approval gate should concentrate attention on the risky step. When it spreads attention across everything, people stop paying it.
Observer (L5)
At L5 the only controls are the activity log and the off switch. Both need design. The log should read like a timeline a manager can scan, with errors at the top and actions grouped by outcome. A raw dump of tool calls helps nobody. The off switch should be easy to find and should leave the system in a known state.
Treat L5 with caution. The CSA framework recommends against it for current deployments. For most business products, L4 with good conditions gives almost all the speed with a fraction of the risk.
Let the dial move
Autonomy should change over time, in both directions.
Earn it with evidence. Start a new agent one level lower than you think it needs. Raise the level for a task when eval pass rates and production error reviews support it. Record that decision the way you would record a change in approval limits for a new staff member.
Let users turn it down. A user who has watched the agent make a mistake should be able to move it back to "ask me first" in one step. The HAX guidelines call this providing global controls. Show the current level in the interface so nobody has to guess.
Keep the dial out of the agent's hands. The CSA framework states that agents should not modify their own autonomy level or disable oversight. That is a product rule as much as a security one.
Questions to ask before you buy or build
Gartner predicts that over 40% of agentic AI projects will be cancelled by the end of 2027, citing rising costs, unclear value and weak risk controls. An explicit autonomy decision per task addresses the third of those head-on.
Many teams will buy agents from vendors instead of building their own. If you are buying an agent, ask the vendor which level each action runs at and who can change it. Then ask to see the approval screen. A vendor who cannot answer has not designed it.
For the broader picture of trust, transparency and testing, read our pillar guide, UX for AI: Designing AI Products People Trust. For the technical side of agents, see how to build an AI agent.
If you are deciding where an agent fits in your operations, our AI product and product design consultation starts by mapping each task to a level and then designs the controls that go with it. You can book a call to walk through one workflow. Teams who want to learn agent UX hands-on can look at the training programmes Marcus teaches.