Researchers demonstrated that encrypted chain-of-thought blocks returned by major LLM APIs (OpenAI, Anthropic, Google) share encryption keys within the same model family. By replaying encrypted reasoning traces into weaker sibling models and jailbreaking them, they could extract the stronger model's hidden reasoning in plaintext. All providers acknowledged the vulnerability and patched it, though Claude Haiku 4.5 was the easiest target.
Background
Major LLM providers have started returning encrypted chain-of-thought reasoning traces to clients as part of their API responses, intended to improve transparency while protecting proprietary reasoning. This vulnerability reveals a critical flaw in how these encrypted blocks are handled across model variants.
- Source
- Simon Willison
- Published
- Aug 12, 2026 at 06:40 AM
- Score
- 7.0 / 10