Anthropic's Frontier Red Team published results showing that frontier AI models (GLM-5.3 and Claude Mythos Preview) can now autonomously develop binary exploitation attacks—achieving full control flow hijacks in 4-6% of trials, a capability earlier models like Claude Opus 4.6 completely lacked. The finding marks a meaningful safety threshold: frontier models are approaching the ability to independently craft low-level cyber attacks.
Background
Anthropic's Frontier Red Team evaluates cutting-edge AI models on security-critical tasks to understand capability boundaries. This benchmark tests whether models can autonomously develop exploitation techniques—a key concern in AI safety research as models approach agentic capabilities.
- Source
- Simon Willison
- Published
- Sep 30, 2026 at 06:20 AM
- Score
- 7.0 / 10