The Double Standard of Closed Labs AI and Aaron Swartz: Training vs Distillation
The Same Act, Two Very Different Outcomes
In 2011, Aaron Swartz wrote a script to download nearly 5 million academic journal articles from JSTOR, with the intention of making them freely available to the public. He believed knowledge should be accessible to everyone, not locked behind paywalls. The government responded with 13 felony charges, carrying up to 35 years in prison. Two years later, at age 26, Aaron Swartz took his own life. In 2024, it was revealed that Anthropic — one of the world's most valuable AI companies — had illegally downloaded and stored 7 million pirated copyrighted books to train its Claude language models. The outcome? A $1.5 billion settlement, approved by a federal judge on July 20, 2026. The largest copyright settlement in U.S. history. No criminal charges. No prison time. Just a business expense.The Double Standard of Closed Labs
| Aaron Swartz | Anthropic | |
|---|---|---|
| What was downloaded | ~5 million academic papers | 7 million copyrighted books |
| Purpose | Free access for everyone | Training proprietary models behind paywalls |
| Commercial gain | None — wanted to give knowledge away | Built a company valued in the hundreds of billions |
| Legal consequence | 13 felony charges, up to 35 years | $1.5 billion settlement (~$3,000/book) |
| Personal cost | His life | A line item on a balance sheet |
The Legal Distinction That Changes Everything
Judge William Alsup of the U.S. District Court for the Northern District of California made a critical ruling:- Training Claude on books constituted fair use — already ruled in previous proceedings
- Downloading and storing 7 million pirated books constituted copyright infringement
What Anthropic Did With That Data
Here's what makes this a double standard, not just a legal case. Anthropic didn't download those 7 million books to share with the world. They downloaded them to train proprietary models locked behind API gates, enterprise contracts, and paywalls. The knowledge extracted from millions of authors' work was absorbed into systems that the public cannot access, inspect, or build upon. Aaron Swartz downloaded academic papers so everyone could read them. Anthropic downloaded books so only paying customers could use the intelligence derived from them. Same method. Opposite intentions.The Same Pattern Across Every Major Lab
Anthropic was simply the first to settle. The exact same allegations are unfolding against every other major LLM provider. Meta — Bartz v. Meta (May 2026)Major publishers (Elsevier, Cengage, Hachette, Macmillan, McGraw Hill) and author Scott Turow sued Meta, alleging it pirated millions of works—from textbooks to novels like The Fifth Season—from the Books3 dataset (Library Genesis) to train its Llama models. The class was certified in 2025, and the case is now proceeding on damages. Meta has not argued that using pirated books constitutes fair use. OpenAI — NYT v. OpenAI (ongoing)
The New York Times sued OpenAI for using millions of articles to train GPT models. In a landmark ruling, the court denied OpenAI's motion to dismiss and explicitly rejected the argument that AI training is "inherently transformative." The case is in discovery, with a trial expected in late 2026 or early 2027. It is widely considered the bellwether case that will define the industry's legal future. Google and xAI
In July 2026, book publishers filed copyright infringement lawsuits against Google over its Gemini AI training. Meanwhile, in December 2025, six authors who opted out of the Anthropic settlement filed individual lawsuits against Anthropic, OpenAI, Google, Meta, and xAI. They allege all five companies copied books from well-known pirate libraries—including LibGen, Z-Library, and OceanofPDF—to train their models. The authors are seeking $150,000 in statutory damages per work per defendant, rejecting the $3,000 settlement as a "tiny fraction" of what the Copyright Act allows. The pattern is identical across the board: closed labs scraped and pirated massive amounts of copyrighted text to train proprietary LLMs locked behind API gates. The only difference is how far along they are in the courtroom.
Distillation Is What Aaron Swartz Would Do Today
This is where the narrative flips. Western AI labs are currently accusing foreign companies of "distillation" — using outputs from proprietary models like Claude or GPT to train competing, more accessible models. They frame it as intellectual property theft. But look at what distillation actually does:- It takes intelligence locked behind closed APIs and paywalls
- It reproduces that capability in models anyone can download, run locally, modify, and share
- It democratizes access to knowledge that closed labs built on pirated data
The Real Hypocrisy
The hypocrisy isn't distillation. The hypocrisy is the closed labs themselves. They built their models by:- Scraping the entire open internet without permission
- Pirating millions of copyrighted books without buying them
- Absorbing the world's knowledge into proprietary systems
What Happens Next
The $1.5 billion settlement covers over 480,000 works. Some authors have opted out, calling the payout insufficient, and plan separate lawsuits. Anthropic has been ordered to destroy the pirated copies in its possession. But the models already learned from that data. The knowledge is baked into the weights. Destroying the books doesn't undo the training. Meanwhile, the legal question of whether AI training on copyrighted material constitutes fair use remains unresolved at the appellate level. Judge Alsup's ruling was a single district court decision, and Anthropic's decision to settle means the case will never reach an appeals court to become binding precedent. The door remains open.The Real Question
Aaron Swartz believed that information wants to be free. He was willing to risk everything for that belief — and ultimately lost his life defending it. The closed AI industry operates on the same principle: data, all data, should be available for training models. But they want the rules to apply selectively. Scraping is fair use when they do it. Distillation is theft when anyone else does it. Distillation isn't the problem. It's what happens when people believe that intelligence should not be locked behind paywalls. It's what Aaron Swartz would have done — taking knowledge hoarded by a few and making it available to everyone. The double standard isn't on the side of those who distill. It's on the side of those who got away with pirating millions of books, built empires from it, and then tried to write the rules so no one else could follow.Source: Reuters — "US judge approves Anthropic's $1.5 billion settlement of copyright lawsuit" (July 20, 2026)
Category:
LLM