US intelligence agencies FBI, NSA, and CISA claim Chinese AI companies have covertly extracted billions of tokens from leading US models like GPT and Gemini to bypass massive R&D costs.
- Chinese firms are allegedly using 'industrial-scale distillation' to siphon capabilities from US AI models.
- Companies like DeepSeek and Moonshot AI are accused of violating terms of service to train their own proprietary models.
- The operation involves sophisticated evasion techniques, including proxy 'transfer stations' and shared premium accounts.
In a sweeping joint advisory issued on September 8, the FBI, National Security Agency (NSA), and the Cybersecurity and Infrastructure Security Agency (CISA) have accused several Chinese AI vendors of conducting industrial-scale theft of proprietary capabilities from US-based frontier AI models. The targets include industry leaders such as OpenAI, Anthropic, Google (Gemini), and SpaceX's xAI (Grok).
The core of the accusation centers on a process known as 'distillation.' While distillation is a legitimate academic practice where a larger 'teacher' model trains a smaller 'student' model to improve efficiency, the US government claims these firms have weaponized the process. By extracting billions of tokens across millions of requests, these companies are effectively stealing the 'reasoning' and 'knowledge' of US models to fuel their own development.
Why This Matters
BozokMedia analysis shows that this is a strategic attempt to erode the US competitive advantage in AI. By utilizing 'stolen' intelligence, Chinese firms can drastically reduce their financial expenditure and shorten development timelines. If a company like DeepSeek can claim a training cost of only $5.6M, it is likely because the heavy lifting of research and data curation was already performed by US companies at a cost of billions.
Model extraction and distillation should be treated as a dedicated security event category, rather than simple API abuse.
The advisory highlights specific culprits: DeepSeek is accused of generating synthetic training data via malicious distillation, while Moonshot AI allegedly plundered Claude Fable 5 data for its Kimi-K3 model and GPT-4o data for Kimi-K2. These efforts were not isolated but were likely conducted with the tacit awareness of the Chinese government.
To evade detection, these firms employed sophisticated tactics. They routed requests through remote cloud providers, used third-party aggregators to hide metadata, and utilized a gray market of proxies known as 'transfer stations' to bypass geographic restrictions imposed by US AI providers.
| Accused Entity | Targeted US Model | Alleged Outcome |
|---|---|---|
| DeepSeek | Various Frontier Models | Synthetic training data generation |
| Moonshot AI | Claude Fable 5 / GPT-4o | Kimi-K2 & Kimi-K3 Development |
| Alibaba / Z.AI | Gemini / Grok / GPT | Reduced compute & research costs |
Frequently Asked Questions
1. Is AI distillation illegal?
Distillation itself is a standard ML technique. However, doing so by violating a company's Terms of Service and bypassing security safeguards to steal proprietary data is considered industrial espionage.
2. How did the Chinese firms avoid being caught?
They used a combination of shared premium accounts, API obfuscation, and 'transfer stations' (proxies) to hide their origin and scale of activity.