Stolen Thoughts 研究:加密 reasoning block 可被跨模型解密,泄露 API key 与密码
Original title: Оказывается, hidden reasoning у frontier-моделей можно было буквально украсть за два API-вызова 😨Очень крутая работа Stolen Thoughts про уязвимость reasoning API у OpenAI, Anthropic и Google.Механика, по сути, гениально простая.Когда reasoning-модель думает, полный Chain-of-Thought пользователю не показывают. Но чтобы продолжать диалог, API возвращает клиенту зашифрованный reasoning block, который потом можно отправить обратно.Проблема оказалась в том, что эти блоки были недостаточно привязаны к конкретной модели, сессии и пользователю.Исследователи брали encrypted reasoning от сильной модели — условно Opus — и передавали его более слабой модели того же провайдера. Слабую модель уже гораздо проще jailbreak'нуть и попросить вывести содержимое reasoning.И она фактически становилась decryption oracle и печатала reasoning сильной модели plaintext'ом.Причем показали это для Anthropic, OpenAI и Google.Самое неприятное даже не кража интеллектуальной собственности модели.Исследователи собрали
The title and summary in the selected language are awaiting translation.
Stolen Thoughts 研究发现 OpenAI、Anthropic 和 Google 的 reasoning API 存在漏洞:加密 reasoning block 未与具体模型、会话和用户充分绑定,把强模型的加密 reasoning 传给同厂商弱模型并越狱后,弱模型会以明文输出强模型的推理内容。
研究揭示加密 reasoning block 可跨模型解密,并给出公开轨迹中泄露密钥的实测数据,对智能体基础设施设计有直接参考价值。
Turns out hidden reasoning in frontier models could literally be stolen with just two API calls 😨
A really cool paper, Stolen Thoughts, about a vulnerability in the reasoning APIs of OpenAI, Anthropic, and Google.
The mechanism is, essentially, brilliantly simple.
When a reasoning model thinks, the full Chain-of-Thought isn’t shown to the user. But to keep the conversation going, the API returns an encrypted reasoning block to the client, which can then be sent back.
The problem turned out to be that these blocks weren’t tightly bound to a specific model, session, and user.
Researchers took encrypted reasoning from a strong model—say, Opus—and passed it to a weaker model from the same provider. The weaker model was already much easier to jailbreak and ask to output the contents of the reasoning.
And it effectively became a decryption oracle, printing the strong model’s reasoning in plaintext.
And they demonstrated this for Anthropic, OpenAI, and Google.
The most unpleasant part isn’t even the theft of the model’s intellectual property.
Researchers collected 6 708 public agent trajectories from GitHub and Hugging Face, with encrypted reasoning blocks still inside them, and reconstructed 315 320 reasoning traces.
They found 704 unique sensitive artifacts: API keys, passwords, access tokens, emails, internal URLs, and other private information.
And some of that data existed only inside the hidden reasoning and never appeared in the model’s visible response at all.
So a developer could look at the log, decide:
“Well, there’s nothing secret here,”
push the trajectory to GitHub—and inside the encrypted blob was a password or API key.
They also did another fun experiment with Kimi-K3: they fed it just the first ~1% of reasoning from Opus 4.8, and even such a small seed was enough to make Kimi’s final answer start drifting toward Opus’s phrasing. So reasoning really is a very powerful hidden state that drives subsequent generation.
And the takeaway for all agent development:
encrypted ≠ safe.
If an opaque blob can be moved into another security context, and the server then decrypts it itself and hands it to the model—that’s essentially a bearer token with a huge amount of hidden context inside.
After responsible disclosure, the providers have already closed the hole: the attacks from the paper no longer reproduce. Reasoning blocks are now more tightly bound to the model and session context.
But the work itself is very good.
We’re increasingly building agent infrastructure around models’ hidden state—reasoning, memory, compaction, checkpoints—and we need to start treating these objects exactly like credentials and secrets, not as some harmless technical fields in an API.
stolen-thoughts.com
Source: AI Coder · Telegram · t.me