Jul 25, 2026
AI

Opus 5 prompt injection tests hit zero browser-agent attacks with safeguards

Anthropic says Claude Opus 5 blocked all browser-agent prompt injection tests when Auto Mode defenses were enabled.

Renata Fuchs

By Renata Fuchs · Policy Reporter

· 3 min read

Opus 5 prompt injection tests hit zero browser-agent attacks with safeguards
Photo: The Decoder

Anthropic says Opus 5 prompt injection defenses stopped every browser-agent attack in a 129-scenario test when the model ran inside Anthropic software with Auto Mode enabled. The result matters for AI agent vendors because prompt injection remains one of the main blockers to letting models act through browsers, files and workplace tools without close human supervision.

Prompt injection is the attack pattern where hostile instructions are placed inside content the model reads, such as hidden text on a webpage, in an attempt to override the model’s operating instructions. For browser agents, that can turn ordinary browsing into a security problem if the model is allowed to take actions after reading compromised content.

Anthropic’s system card says the browser-agent attack success rate was zero percent across 129 scenarios with Auto Mode turned on in products such as Claude Cowork. The company also says Opus 5 is nearly immune to prompt injection in its own software, a claim that depends on both the model and the surrounding protections.

Did Anthropic Opus 5 solve prompt injection?

Anthropic’s reported zero percent result applies to a specific browser-agent setup, not to Opus 5 by itself. With Auto Mode off, the system card reports a 3.7 percent attack success rate for Opus 5, while Sonnet 5 performed better in that configuration at 0.93 percent.

Auto Mode adds two separate defensive layers, according to Anthropic. One layer checks incoming data for concealed instructions before the model processes it. The other stops dangerous actions before they are carried out. Anthropic’s framing is that an attacker must defeat both layers independently for the attack to work.

That distinction is material for companies evaluating agentic AI. A benchmark result tied to protective software is different from a model-level fix, especially if customers are using models in their own stacks or through integrations without the same controls. Anthropic did not present the zero percent browser-agent result as a general guarantee across every deployment pattern.

In a broader prompt-injection evaluation by security firm Gray Swan, Opus 5 also improved on the prior generation but did not reach zero. Anthropic reports that after 15 attempts, the attacker success rate fell to 2.0 percent for Opus 5, down from 5.5 percent for Opus 4.8.

That Gray Swan result put Opus 5 ahead of Mythos 5 at 2.6 percent and Fable 5 at 2.8 percent, according to Anthropic’s chart. The gap is meaningful in a security benchmark, but the remaining 2.0 percent success rate shows that prompt injection is still present outside the specific browser-agent setup protected by Auto Mode.

The broader context is that major AI labs have treated prompt injection as a hard unresolved problem. OpenAI said in December that prompt injection may never be fully solved, according to a report cited in connection with the benchmark discussion. Anthropic’s Opus 5 results suggest progress, but they also underline the practical point for agent builders: the safety boundary is the full product architecture, not only the model name on the API call.

This story draws on original reporting from The Decoder.

More from AI

All AI →