All articles

How Much Autonomy Should an AI Agent Actually Have? What Gartner and Anthropic's Own Data Show

Gartner predicts 40% of enterprises will demote or decommission autonomous AI agents by 2027 over governance failures. Anthropic's own usage telemetry shows why: trust has to be earned action by action, not granted on day one.

4 min read
Two felt puppet characters in an office: one seated at a laptop, the other standing beside the desk with a hand raised in an approval gesture, reviewing a plain clipboard before the first proceeds.

Two data points published within a few months of each other in 2026 tell the same story from opposite ends: giving an AI agent more freedom to act isn't a dial you crank up until it breaks. It's a design decision, and most companies are still making it badly.

Enterprises Are Already Pulling the Plug

In a May 2026 research note, Gartner predicted that by 2027, 40% of enterprises will demote or decommission autonomous AI agents because of governance failures -- not because the underlying models weren't capable enough, but because the governance wrapped around them didn't match how much freedom the agent actually had.

Gartner's framework splits agent autonomy into four levels: agents that only observe (read-only access, outputs visible to one user), agents that advise (recommendations a human still acts on), agents that act with approval (every action requires a human sign-off first), and agents that act autonomously within guardrails, with humans reviewing exceptions and aggregated outcomes rather than individual decisions. The analyst behind the report, Varma, put the core failure mode plainly: enterprises are "treating AI agent governance as binary, either locked down or fully trusted, and that is the root cause of failure." Apply the same blanket controls to a read-only summarization agent and a agent that can send emails or change configurations, and you either strangle the harmless one or leave the risky one under-supervised.

What Actually Happens When You Give an Agent Room to Act

Anthropic's own February 2026 telemetry study of Claude Code usage is a useful real-world counterpoint, because it's not a framework -- it's what actually happened when real users decided, action by action, how much rope to give an agent.

The pattern: newer users (fewer than 50 sessions) used full auto-approve about 20% of the time. By 750 sessions, that climbed past 40%. Trust in an agent, in other words, is earned incrementally and through direct experience, not granted up front. And more autonomy didn't mean less oversight in the way you'd expect -- experienced users who auto-approved more often also interrupted the agent more often (roughly 9% of turns, versus 5% for newer users), suggesting that comfort with autonomy tracks with a sharper sense of exactly when to step back in, not a general loosening of attention.

The broader safety picture in the same data: 80% of tool calls came from agents with at least one safeguard in place (restricted permissions or a human-approval requirement), 73% had some form of human in the loop, and only 0.8% of actions were irreversible, like sending a message that can't be unsent. Between October 2025 and January 2026, the longest agent sessions nearly doubled in duration, success rates on the hardest internal tasks doubled, and the average number of human interventions needed per session still dropped, from 5.4 to 3.3. Agents got more capable and needed less hand-holding at the same time -- but essentially none of the sessions studied were running with zero human touch.

The Real Lesson Isn't "More Autonomy" or "Less" -- It's Matching the Two

Neither data point argues for maximizing autonomy or minimizing it as a blanket policy. Gartner's warning is about mismatched governance: applying Level 4 trust to an agent that should still be at Level 1, or applying Level 1 friction to an agent that's earned more room. Anthropic's data shows what well-matched autonomy looks like in practice -- approval gates that loosen gradually, tied to demonstrated task success, with intervention capability that never fully disappears even at high trust levels.

What This Means Going Forward

The 40% figure from Gartner isn't a prediction that agentic AI fails -- it's a prediction that ungoverned agentic AI, deployed with the same one-size-fits-all trust setting regardless of what the agent is actually doing, fails often enough that a lot of enterprises will end up walking it back. The fix documented in Anthropic's own usage data isn't complicated: approval requirements that scale with the stakes of the action, a human who can always interrupt, and trust that gets extended in proportion to a track record -- not granted on day one and hoped for the best.

Workmate

See what an agent team would do for your business.

Talk to us about Workmate