OpenAI has leveled serious accusations against China-based artificial intelligence company Moonshot AI, alleging a coordinated effort to extract hidden reasoning from its AI models, including ChatGPT. The extracted information was purportedly used to train Moonshot AI's Kimi model through a technique OpenAI terms "adversarial distillation."
Allegations of 'Adversarial Distillation'
According to OpenAI, "adversarial distillation" involves repeatedly crafting specific prompts to compel an AI model to reveal its internal thought process for arriving at an answer. This "protected reasoning," normally withheld from the final output, could potentially allow others to replicate or enhance the model's capabilities without direct access to its core algorithms or training data.
OpenAI clarified that Moonshot AI did not breach its systems, hack encryption, or access stored data. Instead, the alleged method involved manipulating interactions with OpenAI's publicly accessible models to make their internal reasoning visible.
Timeline of the Detected Campaign
The alleged campaign reportedly began on July 1, intensifying later that month. OpenAI stated it detected and subsequently halted a large-scale attempt to extract this hidden reasoning. On July 24 and 25, the company observed a surge of up to 16,000 requests utilizing a particular prompting technique, originating from over 4,000 distinct users. Further investigation revealed more than 15,000 users employing similar prompts.
By July 28, OpenAI successfully disabled the entire campaign, preventing further alleged extraction of its models' internal processes.
Connecting Moonshot AI and the Kimi Model
While OpenAI expressed uncertainty about whether all individuals involved were part of a single group, it believes a core group of participants was associated with Moonshot AI. Notably, Moonshot AI launched its Kimi K3 model on July 16, 2026, which falls within the timeframe of the alleged distillation campaign. If the allegations hold true, this suggests that OpenAI's models may have inadvertently contributed to the development and training of Kimi K3.
OpenAI's Security Response
In response to these events, OpenAI emphasized that defending against AI model distillation requires sophisticated, multi-layered security measures that can adapt to evolving attack techniques. The incident highlights the ongoing challenges in protecting proprietary AI model capabilities and intellectual property in a rapidly advancing technological landscape.