Distillation Technologies Request Access

Home  /  Newsroom  / 

One day apart

On 21 July a judge gave final approval to the largest copyright settlement in US history, against Anthropic, for pirating seven million books. On 22 July the White House accused a Chinese lab of taking Anthropic’s work.

Analysis Sourced

Two things happened in the same week. Taken separately, each is an ordinary news item. Taken together, they describe the position the frontier AI industry has argued itself into.

21 July: the settlement is approved

A federal judge gave final approval to Anthropic’s $1.5 billion settlement with a class of authors and publishers in Bartz v. Anthropic, the largest copyright class-action payout in American legal history. Roughly $3,000 was allocated per work across approximately 500,000 titles. More than 92 percent of eligible rightsholders filed claims. As a condition, Anthropic agreed to destroy the pirated files it had accumulated.

The case was brought by three authors: Andrea Bartz, Charles Graeber and Kirk Wallace Johnson.

22 July: the accusation

The following day, the Director of the White House Office of Science and Technology Policy stated that Moonshot AI had distilled Anthropic’s Fable model to build Kimi K3, and the Treasury Secretary raised the prospect of sanctions. We covered that here.

Within roughly twenty-four hours, the same company was the defendant in the largest case ever brought over training material, and the injured party in a national-security complaint about training material.

What the court actually decided, which matters

It is important to be precise, because the two situations are not the same and the difference is the whole argument.

The court did not find that training an AI model on books was unlawful. It found the opposite: training on the books was fair use. What Anthropic paid $1.5 billion for was acquisition. It had downloaded and retained a central library of roughly seven million pirated books, and holding that library infringed the authors’ rights regardless of what was subsequently done with it.

So the settlement is not authority for the proposition that training on other people’s work is theft. It is authority for the narrower proposition that stealing the copies is theft.

Why distillation is a different legal animal

Distillation claims do not run on copyright at all, for a simple reason: model outputs are not copyrighted. There is no protected work being reproduced when a student model trains on a teacher’s responses.

The claims made against DeepSeek, Moonshot, MiniMax and Alibaba’s Qwen lab therefore rest on breach of terms of service and fraudulent account creation, which are contract and fraud theories, not intellectual property ones. They are much narrower hooks, and they are about how access was obtained rather than about the value of what was learned.

Anyone claiming the book settlement proves distillation is illegal has the law backwards. If anything, the fair-use holding cuts the other way.

The asymmetry that survives

The legal distinction is real. The rhetorical position is harder to defend.

The frontier labs have argued, successfully, that training a model on an enormous corpus of other people’s creative work without permission or payment is fair use, because what the model extracts is patterns rather than expression. That argument won.

The same labs argue that when another laboratory trains on their model’s outputs, it is free-riding, capability theft and a matter for sanctions. OpenAI used the phrase “free-ride on the capabilities developed by OpenAI and other US frontier labs.”

An author whose book sat in that library of seven million could reasonably observe that they had also developed a capability, at cost, which was subsequently used to train a commercial product without their consent, and that when they said so they were told it was fair use.

The strongest defence of the position

There is a coherent answer available, and it deserves stating rather than skipping.

It runs like this. The books were acquired illegitimately, and Anthropic paid for that. Distillation, by contrast, involves accessing a service through a commercial interface subject to an agreement, and then breaching that agreement, frequently while creating thousands of fraudulent accounts to do it at volume. One is a bad acquisition later paid for; the other is an ongoing breach of a contract freely entered into. Consistency does not require treating them identically.

That argument holds. What it does not do is support the language actually used. “Breach of our terms of service” is a much smaller claim than “theft of American AI,” and it does not obviously justify sanctions, Entity List designation, or the involvement of the Treasury.

Where it leaves the argument

The industry has now established, at a cost of $1.5 billion and considerable litigation, that training on other people’s material is lawful and that stealing the copies is not. That is a workable rule. It is also a rule that, applied evenly, makes most of the distillation accusations of 2026 into contract disputes rather than national-security incidents.

Whether the industry wants that rule applied evenly is the open question, and the answer so far appears to be no.