Is It Legal to Train AI on Copyrighted Books? The Law Is a Mess
A $1.5 billion ruling against Anthropic looked like a win for authors — until you read the judge's reasoning.
Published: 2026-08-24 Category: Quick Take Sources: TechCrunch
What's Happening
The AI models behind ChatGPT, Gemini, and Claude are trained on hundreds of millions of books, articles, and papers — most of it without the authors' knowledge or consent. That feels illegal. But the law, as attorney Cathy Gellis told TechCrunch, is "very complex" with "a lot of raw feelings." Last year, Judge William Alsup ordered Anthropic to pay a $1.5 billion settlement to writers whose works were used in training. On the surface, a moral victory for authors. In reality, Alsup ruled that Anthropic's AI training was lawful — the penalty was for pirating books from illegal shadow libraries, not for training on them.
The judge compared an LLM ingesting trillions of words to a writer studying literature: "to turn a hard corner and create something different." Gellis sees this as "generally good news for AI training" because copyright law "hinges on copying, but it doesn't hinge on using the work or experiencing the work, reading the work."
Why It Matters
The core problem is that copyright law hasn't been updated since 1976. Judges are interpreting 50-year-old guidelines to decide questions that will shape the entire AI industry. As Jason Henderson of JWL International put it, "the law is all over the place" because "the AI model has been trained on so much stuff, and the law has not really caught up."
The battleground is fair use — whether training is "transformative" enough to be permissible. Courts weigh the purpose, the amount used, and the impact on the market. Henderson's read of the emerging pattern: "What's tending to win is if you're training on somebody's property because your purpose is to directly compete, then the courts will frown on it. If what you're doing is not going to compete, then the courts are tending to find ways that it will be okay." He points to Thomson Reuters v. Ross Intelligence, where copying content to build a competing AI legal platform drew a lawsuit.
The Catch
The $1.5 billion figure is a distraction. To a company projecting roughly $200 billion in annual revenue by 2028, it's pocket change — and the ruling actually legitimized the core practice of training on copyrighted works. The real signal is that courts are drawing a line between "reading" (training) and "copying" (pirating the source). That's a distinction that favors AI companies far more than authors.
But the legal uncertainty is the real cost. With no legislative update in half a century, every new case is a coin flip that could reshape the industry. Companies are effectively operating on borrowed time, betting that the fair-use interpretation holds while the law catches up.
Bottom Line
The Anthropic ruling was a moral victory for authors and a legal victory for AI companies — both at once. The industry's training practices are increasingly being treated as lawful "reading," with penalties reserved for how the data was obtained, not that it was used. Until Congress updates a 1976 law, the AI copyright question stays a high-stakes gamble decided case by case.
Source: TechCrunch, "Is it legal to train AI models on copyrighted books? It's complicated" (Aug 23, 2026).