Schneier iconSchneierSep 8, 2026 ~6 min source read

How Encrypted 'Reasoning Traces' from LLM APIs Were Decoded by Cross‑Model Replay

Researchers found that encrypted client-side reasoning blobs returned by major LLM providers can be replayed into weaker models within the same provider ecosystem to recover the original, plaintext reasoning. The flaw enables IP extraction, private data leaks, hazardous-reasoning disclosure, and invisible prompt injection.

Share this story

Send the public story page.

Useful takeaways from this story.

Encrypted chain‑of‑thought blobs returned to clients were interoperable across sessions and models, enabling a scalable decryption jailbreak.

Researchers decoded 315,320 publicly scraped reasoning blocks and found 367 PII artifacts plus 182 credentials, showing real-world data leakage.

Attack methods include anti‑distillation circumvention, large‑scale private data extraction, exposure of hazardous internal reasoning, and invisible prompt injection.

# What the research found Leading LLM providers began returning models' step‑by‑step reasoning (chain‑of‑thought) to clients as encrypted blobs rather than storing them server‑side. The idea was to protect intellectual property and reduce information leakage. The research described a structural weakness: those encrypted blobs were compatible and interchangeable across sessions, users, and different models within a provider's ecosystem.

# The exploit in plain terms

# Four concrete attack vectors demonstrated

  • Anti‑distillation circumvention: The technique sidesteps protections meant to prevent extraction of a model's internal logic or reasoning steps, enabling an attacker to reconstruct proprietary reasoning pipelines. The paper demonstrates this across Anthropic, OpenAI, and Google.
  • Large‑scale private data extraction: Developers sometimes publish session logs that include these encrypted blocks. By scraping public repositories, researchers decoded 315,320 reasoning blocks and recovered 367 pieces of personally identifiable information and 182 credentials embedded inside.
  • Hazardous‑information leakage: Even when a model's user‑facing answer rejects a dangerous request, the model's hidden reasoning can contain the hazardous content. The exploit makes that hidden content visible.
  • Invisible prompt injection: Attackers can place malicious payloads entirely inside encrypted blobs. Those blobs can later be replayed into other models to inject instructions without the payload ever appearing in plaintext in public logs, enabling covert poisoning of agentic deployments.

# Scope and affected parties The research tested major providers and reported successful extraction across Anthropic, OpenAI, and Google. The study combined prior work on encrypted reasoning blobs with new cross‑model replay techniques. After responsible disclosure, the providers implemented server‑side fixes, and related reporting indicates mitigation activity.

# Practical implications for developers and operators If your application logs or shares session artifacts, encrypted reasoning blobs in those logs can contain recoverable secrets or internal instructions. Publicly posted session dumps, backups, or developer troubleshooting files should be audited for such blobs and treated as potential data leaks. Relying on client‑side opaque tokens for protecting sensitive intermediate model state is risky when those tokens are interoperable across models.

# High‑level mitigations mentioned The authors proposed cryptographic and system‑level fixes to secure client‑side reasoning. Providers have since deployed server‑side mitigations according to public reports. Practical steps for teams include revoking or removing published session logs, rotating exposed credentials, and auditing agent rollouts for replayable encrypted blocks.

# Bottom line An architectural choice—returning encrypted reasoning to clients as interoperable blobs—created an unexpected cross‑model replay attack path. That path exposed proprietary reasoning, private data, and new prompt‑injection vectors. Fixes exist and have been applied, but the breach highlights the risk of treating opaque client tokens as sufficient protection for sensitive model internals.

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app