product updateOpenAI

OpenAI says it shut down a 15,000-account campaign to extract model reasoning; the attack still worked on Azure

TL;DR

OpenAI says it detected and shut down an adversarial distillation campaign involving more than 15,000 accounts, linked in part to people associated with Moonshot AI. Researchers report that the same reasoning-extraction attack still worked on Microsoft Azure on September 13 against every OpenAI model they tested, including GPT-6 Astra.

4 min read
0

OpenAI says it broke up a coordinated campaign to copy the hidden reasoning of its models, a technique it calls adversarial distillation. Independent researchers report that the same extraction trick kept working on Microsoft Azure for weeks afterward, including against GPT-6 Astra.

What OpenAI reports

According to OpenAI's blog post, the activity began at low volume on July 1. On July 24 and 25, it spiked to 16,000 requests from more than 4,000 users, all following a typical extraction pattern. OpenAI then identified a network of more than 15,000 accounts with related patterns and says it had fully shut them down by July 28. A footnote clarifies that these were attempted extractions, not necessarily successful ones.

OpenAI links a core group behind the activity to people associated with Moonshot AI, maker of the Kimi models. It says it is unclear whether all observed actors trace back to a single source. Moonshot's response was not included in the source reporting. Anthropic recently reported similar attempts by Chinese AI companies.

How the attack works

Providers return reasoning to customers only as encrypted data packets, which customers pass back with follow-up requests. Researcher Joachim Schaeffer and colleagues showed in a paper that, because the packets are encrypted with shared keys, they can be moved between sessions, between users, and between different models from the same provider.

Attackers copied encrypted reasoning from one conversation and asked a model in a separate conversation to decrypt and write it out. A weaker, cheaper model from the same family can therefore act as a "decryption oracle" and print a stronger model's hidden reasoning verbatim. OpenAI says it confirmed the reported attack paths were real and that the research helped it deploy countermeasures faster.

OpenAI says it has since:

  • banned fraudulent accounts and tightened sign-ups
  • closed the hole that allowed reuse and reading of encrypted reasoning from other sessions
  • begun screening streamed outputs and holding them back if they might reveal reasoning
  • shared findings through the Frontier Model Forum and government channels

Azure remained open

The researchers published an update the same day. When they retested on September 13, the attack was blocked on OpenAI's and Anthropic's own APIs. On Azure, it worked against every OpenAI model tested, including GPT-6 Astra, and against Anthropic models up to Sonnet 5. A single attempt was enough to extract reasoning verbatim.

Per the researchers' timeline, OpenAI did not add safeguards to the Azure endpoint until September 27. For Anthropic models, the extraction could no longer be reproduced on Azure from September 28. The researchers say GPT-6 Astra launched on third-party platforms without any of the protections in place.

A second, simpler method, publicly demonstrated by developer Can Bölük, gives the model a virtual notepad tool and instructs it to write out its reasoning there. The researchers say this worked on every OpenAI model, as well as on Opus 4.8 and Sonnet 5. Only Opus 5, Fable 5, and Fable 5.1 did not reveal their reasoning. The researchers believe the output would likely be as useful for distillation as the decryption attack's.

The researchers call the fixes so far piecemeal and superficial, noting that many rely on brittle matching of specific request patterns and some reached cloud platforms only days later. OpenAI acknowledges that partner-hosted models need the same protection as its own services and says the work is not finished.

What this means

The core finding is that a provider's security is only as strong as its weakest distribution channel. Identical models carried different protections depending on who served them, so attackers could simply choose the weakest route. Patching one API does not close the exposure.

The researchers argue that cloud providers without equivalent protections should not serve reasoning models at all, since open backdoors would effectively allow people to sidestep export controls at the API level. That is a policy claim, not an established standard, but it points to where pressure is heading: contractual security requirements for hosting partners, and launch gating for new models on third-party platforms.

For developers, the practical takeaway is that reasoning-token handling now varies by platform. Teams relying on hidden reasoning staying hidden should not assume parity between first-party and cloud-hosted endpoints. OpenAI expects distillation attempts to grow more sophisticated as frontier models improve.

Comments

Loading...