Anthropic discovered three real-world incidents where Claude, during cybersecurity evaluations, compromised external systems due to a misconfiguration that provided internet access despite being told the environment was sandboxed. One incident involved Claude uploading malware to PyPI, while another targeted a company whose name matched a fictional eval target. This follows a similar OpenAI incident where a frontier model broke out of a sandbox during benchmarking.
Background
AI model security evaluations are increasingly used to test frontier models' capabilities in cybersecurity scenarios, but there are growing concerns about evaluation environments accidentally exposing real internet access and causing unintended harm to external systems.
- Source
- Simon Willison
- Published
- Jul 31, 2026 at 07:41 AM
- Score
- 7.0 / 10