The NSA, FBI, and CISA accuse six China-based AI companies of running industrial-scale knowledge-distillation campaigns against U.S. frontier AI models. The September 8, 2026 advisory names DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI. Separately, Anthropic says Alibaba-linked operators generated more than 151 million Claude exchanges between May and July 2026, while it attributed more than 23 million exchanges to Moonshot AI and more than 12.1 million to DeepSeek.
Those are serious allegations, but “distillation” is not another word for theft. The technical method is legitimate; the dispute is about who accessed which models, under what authorization, through which accounts and intermediaries, and at what scale. The allegations concern outputs, reasoning traces, and capabilities obtained through model interfaces—not an established transfer of source code or model weights.
What the U.S. advisory says
The joint advisory says organized campaigns have operated since at least late 2024 and targeted models including Claude, GPT, Gemini, and Grok. It describes alleged use of native APIs, cloud providers, third-party aggregators, and gray-market “transfer stations” to reach models or bypass geographic restrictions.
The alleged access techniques include fraudulent accounts, shared premium subscriptions, proxies, disposable email addresses, virtual-card payments, metadata sanitization, and automated switching between access routes. The companies named in the advisory are distinct from Anthropic’s separate reference to seven China-based labs; the two lists should not be treated as identical.
The central allegation is that these access networks collected large volumes of model outputs to help competing systems learn capabilities such as reasoning, coding, software engineering, mathematics, reinforcement learning, and agentic tasks. Naming a company in the advisory is an official U.S. government allegation, not a judicial finding that the company violated a specific law.
Distillation versus unauthorized extraction
In ordinary model distillation, a more capable teacher model generates answers to selected prompts. A separate student model uses those outputs as training material to learn particular behaviors or capabilities, often with the goal of producing a smaller or more efficient system. The technique does not inherently require access to source code or model weights.
The disagreement begins at the authorization boundary. Anthropic and the U.S. agencies allege that some operations used fraudulent identities, circumvented geographic restrictions, routed requests through intermediaries, or collected restricted reasoning capabilities at industrial scale. Those circumstances—not the existence of teacher-student training itself—are what turn the case into a security, contractual, privacy, and policy dispute.
| Dimension | Ordinary distillation | Conduct alleged by Anthropic and U.S. agencies |
| Purpose | Use a teacher model’s outputs to train a student model | Extract capabilities from third-party frontier models for competing systems |
| Access | Provider-permitted access, licensing, research, or a company’s own models | Alleged use of fraudulent accounts, proxies, aggregators, and transfer stations |
| Authorization | Access and data use follow the provider’s permission | Alleged circumvention of geographic restrictions and possible terms-of-service violations |
| Material obtained | Outputs or training signals; source code and weights are not inherently required | Outputs, reasoning traces, proprietary functionality, and synthetic training data are at issue |
| Legal meaning | The technique itself is not automatically unlawful | The allegations are not a court finding of theft or infringement |
The distinction matters because calling every form of distillation “theft” would erase the difference between a standard training strategy and alleged abuse of an access system.
The Claude exchange figures
Anthropic’s September 2026 threat-intelligence report attributes the following Claude activity to operators linked to three China-based companies. The figures are attributed assessments, not independently adjudicated measurements.
| Entity | Attributed Claude activity | Measurement window | Additional condition |
| Alibaba-linked operators | More than 151 million exchanges | May–July 2026 | Activity peaked at nearly 3 million exchanges per day and involved more than 3,500 allegedly fraudulent accounts |
| Moonshot AI | More than 23 million exchanges | May–July 2026 | Anthropic also described nearly 300,000 customer requests routed to Claude during one 10-day period |
| DeepSeek | More than 12.1 million exchanges | 14 days in July 2026 | Anthropic attributed the activity to DeepSeek |
The scale is the striking part. A single account or an isolated experiment would raise different questions from millions of exchanges spread across thousands of allegedly fraudulent accounts. But the size of a number does not resolve attribution or legal responsibility by itself; it describes the activity Anthropic says it identified.
The customer-privacy question
Anthropic says some Moonshot AI/Kimi and DeepSeek requests were allegedly routed to Claude while users believed they were interacting with the Chinese companies’ models. The allegation raises a separate issue from model training: whether people’s prompts were sent to another provider and whether customers were clearly informed.
Anthropic also said some exchanges included sensitive information from individual users, multinational companies, and state-affiliated actors. That makes the alleged routing more consequential than a simple benchmark comparison. A user prompt can contain confidential business material, personal information, or internal code, so the identity of the model processing it matters.
The customer-notification question remains part of the allegation. It should not be converted into a definitive claim that every affected user was kept unaware.
What providers may do next
CISA recommends that model providers detect anomalous behavior, share intelligence across providers and cloud platforms, and alter responses for high-confidence malicious-distillation attempts. The advisory specifically describes switching suspected malicious requests to lower-fidelity or downgraded outputs while treating legitimate researchers and evaluators differently.
That is a targeted provider recommendation, not an established blanket policy for Chinese users and not evidence that every user in China will receive degraded answers. The practical challenge is classification: providers must distinguish a coordinated extraction operation from legitimate evaluation, research, or ordinary use without punishing the broader user population for the behavior of suspected accounts.
What the allegations establish—and what they do not
The September 8 advisory establishes that the NSA, FBI, and CISA publicly accused six named China-based AI companies of industrial-scale distillation campaigns. Anthropic’s September report establishes that the company attributes large volumes of Claude exchanges to Alibaba-linked operators, Moonshot AI, and DeepSeek and says it detected and disrupted the activity it describes.
The allegations do not amount to a court ruling that the companies stole U.S. AI models. The public case described here concerns outputs, reasoning traces, access methods, and proprietary capabilities obtained through model interfaces. It does not establish a transfer of source code or model weights.
For readers, the important boundary is simple: distillation is a normal machine-learning technique, but large-scale access built on allegedly fraudulent accounts or bypassed restrictions is a very different proposition. The next consequences will depend on how providers enforce access rules and whether the allegations move beyond company and government assessments into formal legal findings.