A security researcher demonstrates a 60-80% prompt injection attack success rate against Claude Code Opus 5 Auto Mode, contradicting Anthropic's commissioned evaluation that reported 0.00% success. The attack chains multiple steps—forcing curl over WebFetch, exploiting a poisoned Python module—to achieve code execution, raising concerns about the safety of Auto Mode as the default.
Background
Auto Mode was introduced as the default for Claude Code in mid-August 2026, replacing human approval prompts with an AI safety classifier. Prompt injection remains a critical vulnerability in LLM-based agents, especially when they have tool access and code execution capabilities.
- Source
- Lobsters
- Published
- Aug 30, 2026 at 01:36 PM
- Score
- 8.0 / 10