The Same Act, Two Very Different Outcomes

In 2011, Aaron Swartz wrote a script to download nearly 5 million academic journal articles from JSTOR, with the intention of making them freely available to the public. He believed knowledge should be accessible to everyone, not locked behind paywalls. The government responded with 13 felony charges, carrying up to 35 years in prison. Two years later, at age 26, Aaron Swartz took his own life. In 2024, it was revealed that Anthropic — one of the world's most valuable AI companies — had illegally downloaded and stored 7 million pirated copyrighted books to train its Claude language models. The outcome? A $1.5 billion settlement, approved by a federal judge on July 20, 2026. The largest copyright settlement in U.S. history. No criminal charges. No prison time. Just a business expense.

The Double Standard of Closed Labs

Aaron Swartz Anthropic
What was downloaded ~5 million academic papers 7 million copyrighted books
Purpose Free access for everyone Training proprietary models behind paywalls
Commercial gain None — wanted to give knowledge away Built a company valued in the hundreds of billions
Legal consequence 13 felony charges, up to 35 years $1.5 billion settlement (~$3,000/book)
Personal cost His life A line item on a balance sheet
The scale of Anthropic's infringement was far more extensive than Swartz's. The commercial benefit was immeasurable. And the punishment was a fraction of what the company earned from the very data it pirated.

The Legal Distinction That Changes Everything

Judge William Alsup of the U.S. District Court for the Northern District of California made a critical ruling:
  • Training Claude on books constituted fair use — already ruled in previous proceedings
  • Downloading and storing 7 million pirated books constituted copyright infringement
The problem wasn't that Anthropic learned from the books. The problem was how they got them. They didn't buy the books. They pirated them. This created a bizarre reality: the act of learning was legal, but the act of obtaining the material was not — and only the latter carried a price tag.

What Anthropic Did With That Data

Here's what makes this a double standard, not just a legal case. Anthropic didn't download those 7 million books to share with the world. They downloaded them to train proprietary models locked behind API gates, enterprise contracts, and paywalls. The knowledge extracted from millions of authors' work was absorbed into systems that the public cannot access, inspect, or build upon. Aaron Swartz downloaded academic papers so everyone could read them. Anthropic downloaded books so only paying customers could use the intelligence derived from them. Same method. Opposite intentions.

The Same Pattern Across Every Major Lab

Anthropic was simply the first to settle. The exact same allegations are unfolding against every other major LLM provider. Meta — Bartz v. Meta (May 2026)
Major publishers (Elsevier, Cengage, Hachette, Macmillan, McGraw Hill) and author Scott Turow sued Meta, alleging it pirated millions of works—from textbooks to novels like The Fifth Season—from the Books3 dataset (Library Genesis) to train its Llama models. The class was certified in 2025, and the case is now proceeding on damages. Meta has not argued that using pirated books constitutes fair use. OpenAI — NYT v. OpenAI (ongoing)
The New York Times sued OpenAI for using millions of articles to train GPT models. In a landmark ruling, the court denied OpenAI's motion to dismiss and explicitly rejected the argument that AI training is "inherently transformative." The case is in discovery, with a trial expected in late 2026 or early 2027. It is widely considered the bellwether case that will define the industry's legal future. Google and xAI
In July 2026, book publishers filed copyright infringement lawsuits against Google over its Gemini AI training. Meanwhile, in December 2025, six authors who opted out of the Anthropic settlement filed individual lawsuits against Anthropic, OpenAI, Google, Meta, and xAI. They allege all five companies copied books from well-known pirate libraries—including LibGen, Z-Library, and OceanofPDF—to train their models. The authors are seeking $150,000 in statutory damages per work per defendant, rejecting the $3,000 settlement as a "tiny fraction" of what the Copyright Act allows. The pattern is identical across the board: closed labs scraped and pirated massive amounts of copyrighted text to train proprietary LLMs locked behind API gates. The only difference is how far along they are in the courtroom.

Distillation Is What Aaron Swartz Would Do Today

This is where the narrative flips. Western AI labs are currently accusing foreign companies of "distillation" — using outputs from proprietary models like Claude or GPT to train competing, more accessible models. They frame it as intellectual property theft. But look at what distillation actually does:
  • It takes intelligence locked behind closed APIs and paywalls
  • It reproduces that capability in models anyone can download, run locally, modify, and share
  • It democratizes access to knowledge that closed labs built on pirated data
That is exactly what Aaron Swartz would have done if he were alive today. Swartz didn't create the academic papers — he liberated them from behind paywalls so everyone could benefit. Distillation doesn't create model intelligence from scratch — it liberates it from behind API gates so everyone could benefit. Both acts take something hoarded by a few and make it available to all. Both challenge the idea that knowledge should be locked up for profit. The only difference is the medium: Swartz fought paywalled journals. Today's distillers fight paywalled models.

The Real Hypocrisy

The hypocrisy isn't distillation. The hypocrisy is the closed labs themselves. They built their models by:
  1. Scraping the entire open internet without permission
  2. Pirating millions of copyrighted books without buying them
  3. Absorbing the world's knowledge into proprietary systems
And now they want the rules to apply selectively: fair use for their training, theft for everyone else's distillation. If the argument is that the open internet is fair game for training, then model outputs — the distilled essence of that same internet — should be too. You can't claim the entire world's knowledge belongs to you while denying others the right to learn from what your models produce.

What Happens Next

The $1.5 billion settlement covers over 480,000 works. Some authors have opted out, calling the payout insufficient, and plan separate lawsuits. Anthropic has been ordered to destroy the pirated copies in its possession. But the models already learned from that data. The knowledge is baked into the weights. Destroying the books doesn't undo the training. Meanwhile, the legal question of whether AI training on copyrighted material constitutes fair use remains unresolved at the appellate level. Judge Alsup's ruling was a single district court decision, and Anthropic's decision to settle means the case will never reach an appeals court to become binding precedent. The door remains open.

The Real Question

Aaron Swartz believed that information wants to be free. He was willing to risk everything for that belief — and ultimately lost his life defending it. The closed AI industry operates on the same principle: data, all data, should be available for training models. But they want the rules to apply selectively. Scraping is fair use when they do it. Distillation is theft when anyone else does it. Distillation isn't the problem. It's what happens when people believe that intelligence should not be locked behind paywalls. It's what Aaron Swartz would have done — taking knowledge hoarded by a few and making it available to everyone. The double standard isn't on the side of those who distill. It's on the side of those who got away with pirating millions of books, built empires from it, and then tried to write the rules so no one else could follow.
Source: Reuters — "US judge approves Anthropic's $1.5 billion settlement of copyright lawsuit" (July 20, 2026)