Encrypted AI reasoning extraction attacks
- Researchers published August 16 that encrypted reasoning blocks at Anthropic, OpenAI, and Google shared a single key per model family, letting attackers replay frontier reasoning into cheaper siblings and jailbreak them into reading the traces aloud; from real sessions they recovered 62 API keys, 33 passwords, and 24 access tokens.
- All three providers have patched the attacks.
- A finding that Kimi K3 produced outputs similar to reasoning from Claude Opus 4.8 and GPT 5.6 Sol is contested: researchers explicitly state they cannot establish causation, and Anthropic called roughly 16 million exchanges via 24,000 Chinese-attributed accounts output harvesting, not distillation.
Shared encryption keys across model families meant securing a flagship model provided no protection if a cheaper sibling could act as a decryption oracle, and real user credentials were recovered from production sessions before patches were applied.