Stealing Reasoning Traces from Proprietary LLM APIs
Researchers extracted hidden reasoning from frontier LLMs by exploiting shared family-wide encryption keys
“Anthropic, OpenAI, and Google return encrypted chain-of-thought blocks to clients that can be replayed across sessions, users, and models. We take a trace produced by a frontier model, replay it into a weaker sibling, jailbreak the weaker model, and recover the stronger model's hidden reasoning in plaintext”
A new paper reveals that Anthropic, OpenAI, and Google all shared a single encryption key across model families for their encrypted reasoning blocks, allowing researchers to replay a strong model's trace into a weaker sibling and jailbreak it into exposing the raw chain-of-thought. Claude Haiku 4.5 was the most susceptible, exploitable via assistant turn prefilling that was subsequently removed in 4.6 models. All three providers have since patched the vulnerability, but the paper's appendix offers a rare public glimpse into the unpolished internal reasoning of frontier models.