The legal status of training AI on copyrighted books without permission remains unresolved.
Why it matters: The answer could shape how developers build and pay for generative AI systems. It also affects publishers and authors whose books are used in training datasets.
- The dispute centers on whether training AI on copyrighted books without permission qualifies as fair use.
- In Anthropic litigation, a federal judge said copying books from pirate libraries was unlawful but treated the training use as fair use.
- Anthropic agreed to pay $1.5 billion to resolve authors' claims tied to pirated books used in training, AP reported.
- The U.S. Copyright Office says its planned Part 3 report will focus on AI training and copyrighted works.
The core legal question is still open: whether training AI models on copyrighted books without permission is lawful under fair use. That issue has become a central fight in copyright cases involving model developers, authors and publishers.
Recent Anthropic litigation illustrates the split. In a federal ruling, the judge found that copying books from pirate libraries was unlawful, while also concluding that the training use itself could be fair use. AP reported that Anthropic later agreed to pay $1.5 billion to resolve authors' claims tied to those pirated books.
The U.S. Copyright Office is still examining the issue. Its AI page says the office has spent years studying copyright and AI, and that its planned Part 3 report will address the legal implications of training AI models on copyrighted works.
Existing copyright precedent gives the debate some direction, but not a clean answer for generative AI. In Authors Guild v. Google, the Second Circuit held that scanning books for search could be fair use because it served a different purpose from reading the books themselves. But that case did not decide the legality of training modern AI models on books at scale.
By the numbers
- $1.5 billion - Anthropic's reported settlement with authors over pirated books used in training.
- Part 3 - The Copyright Office's planned report section focused on AI training and copyrighted works.
Yes, but: The existing precedents cited here address book scanning and authorship questions, not a final court ruling on AI training with copyrighted books.
What's next: The Copyright Office's planned Part 3 report is expected to address AI training and copyrighted works.