Stealing Reasoning Traces from Proprietary LLM APIs

Research official 2 src. ~1 min

Researchers show that encrypted reasoning-trace blocks returned by major LLM APIs are interchangeable across sessions, users, and even different models, letting attackers inject them into weaker models to force verbatim disclosure of the plaintext reasoning.

Why it matters

Documents four attack vectors affecting Anthropic, OpenAI, and Google, recovering 367 PII artifacts and 182 credentials from over 315,000 scraped reasoning blocks; proposes cryptographic and system-level mitigations.

Importance: 2/5

Security research finding below the 100-upvote bump threshold; kept at default notable-update level.

Sources