Prominent prompt injection researcher Johann Rehberger discovered an attack against Anthropic's Claude Code auto mode that works 80% of the time, tricking the agent into executing malicious code from a disguised archive. Auto mode paradoxically blocked Claude's attempt to clean up the malware after detecting it, revealing a critical flaw in the safety mechanism itself. Simon Willison emphasizes that sandboxed environments are essential for running unattended AI agents safely.
Background
Anthropic recently made Claude Code's auto mode the default, claiming strong protection against prompt injection attacks. This research by Johann Rehberger challenges those claims by demonstrating a real-world exploit.
- Source
- Simon Willison
- Published
- Aug 28, 2026 at 06:50 AM
- Score
- 8.0 / 10