Full Breakdown
OpenAI Disrupts Large-Scale “Adversarial Distillation” Campaign Tied to China’s Moonshot AI
By Drooid · · How we work
Core Event: Coordinated Extraction of Protected Model Reasoning
OpenAI announced it had detected and shut down a month-long campaign that sought to pull hidden “protected reasoning” from its frontier AI models. The operation began on July 1, 2026, peaked on July 24-25 with more than 4,000 accounts generating roughly 16,000 matching requests, and was halted by July 28. OpenAI attributes a core cluster of the activity to individuals linked to Moonshot AI, the Beijing-based developer of the Kimi series.
Background & Context
The incident follows earlier accusations by Anthropic that Moonshot AI and other Chinese labs were covertly using Anthropic’s Claude model to train their own systems. A U.S. Cybersecurity and Infrastructure Security Agency (CISA) listing also named Moonshot as a suspected extractor of American-built models, highlighting growing tension over model-theft allegations.
Timeline
- July 1, 2026 – Low-volume probing begins.
- July 24-25, 2026 – Spike to ~16,000 extraction attempts from >4,000 users.
- July 27, 2026 – Security firm Mindgard notifies Moonshot of separate safety-control bypasses in its Kimi models.
- July 28, 2026 – OpenAI reports the campaign fully disrupted and related activity identified across >15,000 users.
- Sept 30, 2026 – OpenAI details the attack in a blog post and shares findings with the Frontier Model Forum and government channels.
Data & Statistics
- 16,000 attempted extraction requests over two days.
- 4,000+ distinct user accounts involved in the peak.
- 15,000+ accounts linked to related prompt-pattern activity across the month.
- The technique copied encrypted chain-of-thought output from one conversation and prompted a separate model instance to decode it, without breaching encryption or accessing stored user conversations.
Official Statements & Responses
- OpenAI called the activity “adversarial distillation,” warning that extracting protected reasoning could let competitors reproduce advanced capabilities without the original safety safeguards. The company banned the offending accounts, tightened sign-up verification, added monitoring for distillation patterns, and closed a “replay pathway” that allowed encrypted reasoning to be reused. Findings were shared with other AI developers and U.S. government programs.
- Caroline Zier, head of OpenAI’s strategic national security policy program, said the concern centers on terms-of-service violations rather than open-weight models and reaffirmed support for an open-weight ecosystem while strengthening defenses.
- Michael Kratsios, technology policy advisor to the White House, noted that Moonshot’s use of Nvidia Blackwell chips—subject to U.S. export restrictions—adds a national-security dimension to model-theft.
- CISA has listed Moonshot AI among organizations it believes are extracting data and reasoning from American-built models.
Conflicting Reports & Gaps
OpenAI said it could not confirm that every participant was linked to Moonshot AI and provided no technical evidence publicly, citing security considerations.
What’s Next
OpenAI will expand monitoring for distillation patterns, tighten API verification, and work with industry and government partners on defenses such as “preserved thinking” mechanisms. The U.S. administration’s framing of model-theft as a national-security issue suggests possible future export-control or regulatory actions affecting cross-border AI development.
