Anthropic makes Claude Code's Auto Mode the default from August 14

AI AgentsDeveloper Tools
Editorial image for Claude Code Will Auto-Approve Actions by Default From Aug 14

The Core · TL;DR

  • Anthropic will make Claude Code's Auto Mode the default for Pro, Max, and Team plans starting August 14, 2026; enterprise users must opt in separately.
  • Testing with 1,053 paid users found Auto Mode caught 89% of dangerous commands versus 13.6% for human reviewers, who approved 97% of prompts regardless of content.
  • An independent Trajectory Labs audit found zero successful prompt injections across 720 attack attempts against Auto Mode; a separate OpenAI comparison figure appears to be misreported.
  • New hard deny rules and prompt injection screening aim to block data exfiltration, and Anthropic doesn't charge for the safety classifier's token usage.

Starting August 14, 2026, Claude Code will stop asking permission before every action for users on Anthropic's Pro, Max, and Team plans. Auto Mode, first tested back in March, becomes the standard experience rather than an opt-in feature.

The shift inverts how coding agents have typically been governed. Instead of pausing at each step for a human to click "approve," Claude Code will now run continuously and only stop when it flags a move as dangerous, irreversible, or directed outside the user's own environment.

Anthropic's justification rests on a blunt observation about human behavior: in practice, users approved 97% of permission prompts regardless of content. That pattern suggests manual review had become a rubber stamp rather than a genuine safeguard.

The numbers behind the switch

A controlled study of 1,053 paid testers found human reviewers caught just 13.6% of dangerous commands, while Auto Mode's own classifier caught 89%. Different outlets reported these figures in reversed order, but they describe the same result.

An independent audit by Trajectory Labs pushed further on the security question, running 72 indirect prompt injection scenarios ten times apiece, 720 attempts in total, against Claude's Fable 5, Opus 5, and Sonnet 5 models in Auto Mode. None succeeded.

That claim sits alongside a murkier comparison point. The Decoder reported that OpenAI's GPT-5.6 Sol let 5.83% of attacks through in Codex Auto-Review mode, but other reporting attributes a 0.05% failure rate to GPT-5.6 Sol against direct attacks, with a separate 19.03% figure apparently describing Claude Code running in Full Access mode rather than OpenAI's model. The comparative benchmark appears to have been garbled somewhere in the reporting chain, and readers should treat any OpenAI-versus-Anthropic injection-resistance comparison from this period with caution.

What changes for developers

Anthropic pairs the default switch with new guardrails: prompt injection screening and customizable hard deny rules meant to block data exfiltration attempts specifically. The company also isn't charging for tokens its safety classifier consumes while making these judgment calls.

The internal dogfooding numbers are notable. During one extended session, Auto Mode reportedly killed around 2,000 processes that would have interfered with live GPU training runs at Anthropic. Teams using it also generated roughly 25% more pull requests than those working under manual approval.

Boris Cherny, who leads Claude Code at Anthropic, said his team now uses Auto Mode exclusively and has no interest in returning to permission prompts. Enterprise customers, notably, are exempt from the automatic switch and must still opt in.

WK

WAKIB Editorial Team

This review was prepared and summarized by the WAKIB AI intelligence engine and vetted by our editorial board for accuracy and reliability.

Subscribe to Newsletter

Get a weekly summary of the most promising AI research and tools delivered to your inbox.

Telegram Channel

Join our active community on Telegram for real-time tracking of AI models and trends.

Join us on Telegram

More from Tools

View all in Tools