The Hallway Track
Research Findings

Breaking Claude Code Opus 5 Auto Mode

Simon Willison · Aug 27, 2026 · Research Findings

Prompt injection attack bypasses Claude Code Auto Mode with 80% success rate

“The safety mechanism itself can become part of the failure. The classifier allowed the creation of the malware process, but then it blocked the command intended to stop it!”

Security researcher Johann Rehberger demonstrated an 80% success rate attack against Claude Code's Auto Mode — Anthropic's recently-defaulted safety mechanism for coding agents — by exploiting Python import shadowing via a malicious zip archive. More alarmingly, Auto Mode blocked Claude's own cleanup commands after it detected the compromise, inverting the safety guarantee. The finding reinforces that sandboxed execution environments remain essential for unattended AI coding agents.

prompt-injection claude-code security ai-agents sandboxing anthropic

Watch / read the original source →