Skip to content
News

Anthropic Details AI Distillation Attacks From China Labs

Anthropic has published a report alleging persistent distillation attacks by China-based AI companies, activity it says has escalated in recent months as competition across the sector has intensified.

According to the report, released on Thursday, unauthorised labs have developed increasingly sophisticated methods to circumvent Anthropic’s defences and harvest the capabilities of US frontier models. The campaigns targeted some of Claude’s most valuable capabilities, including agentic behaviour and tool use, coding and data analysis, and logical reasoning.

The company had previously raised concerns about distillation attacks in February, naming specific labs at the time. OpenAI has reported comparable activity, which it attributed to DeepSeek. The campaigns described in the new report are both larger and more aggressive, with Anthropic observing nearly 200 million exchanges linked to distillation attacks across five separate campaigns.

How distillation attacks work

Broadly, distillation attacks focus on extracting the chain of thought from a model’s response to various queries. That chain of thought can then be used to train a smaller model on general reasoning ability through supervised fine-tuning.

Anthropic typically does not make its models’ internal chain of thought available to users, instead displaying summarised thinking blocks that provide a general overview. The distillation campaigns, however, identified specific techniques that could trick the model into revealing its thinking traces directly. In one instance, an attacker framed a query as a translation request, instructing the model to translate its previous working memory into katakana-only Japanese.

Campaigns linked to Alibaba and Moonshot AI

The bulk of the attempts came from a campaign attributed to Alibaba, described as the largest wholesale distillation effort Anthropic has observed. The company recorded 151 million exchanges between May and July 2026, peaking at nearly three million exchanges per day. The exchanges were spread across 3,500 accounts, but because they shared a single fixed prompt used to extract the chain of thought, Anthropic attributed them to a single effort to produce training material for Alibaba’s Qwen family of models.

A separate campaign from Moonshot AI, the maker of Kimi, appeared to route requests directly from the Chinese military. One request asked Claude to assess a cache of closed-circuit surveillance footage to determine whether the subject was behaving abnormally. Over a 10-day period, nearly 300,000 requests were routed to Claude through a network of 5,000 accounts, primarily targeting the company’s Opus model.

Alongside Alibaba and Moonshot AI, the report also implicated DeepSeek among the China-based companies tied to the identified campaigns.

Source

The UK tech briefing

Smartphones, AI, computing and deals — the essential stories without the noise.

Mailing provider can be connected when your UK list is ready.

Shop on Amazon UK — Discover deals Shop on Amazon UK — Discover deals