Why training AI on copyrighted books lands in legal gray zone

EthicsLLMs
Is it legal to train AI models on copyrighted books? It’s complicated | TechCrunch

techcrunch.com · News coverage photograph, editorial use approved

The Core · TL;DR

  • Fair use, a 1976-era copyright doctrine, is being applied to AI training on hundreds of millions of books and articles with mixed results in court.
  • Judge Alsup ruled Anthropic's AI training was lawful fair use, but the company still owes $1.5 billion for pirating books from shadow libraries.
  • Judge Bibas ruled against fair use in a separate case where AI trained on Reuters content was used to build a directly competing product.
  • Courts appear to be distinguishing between how training data was acquired and whether the resulting AI product competes with the original source.

Every major chatbot, from ChatGPT to Gemini to Claude, learns from datasets stitched together out of hundreds of millions of books, articles, academic papers, and other scraped internet text. Whether that practice is legal depends on a doctrine that predates the internet by decades, and courts are still working out how it applies.

That doctrine is fair use, the copyright carve-out that lets someone use protected material without permission for purposes like criticism, commentary, education, or parody. It was written into US law in 1976, the same year the statute itself was last substantially updated, long before anyone imagined a model ingesting a library's worth of text to predict the next word in a sentence.

Recent rulings show how unsettled the question remains. Judge William Alsup found that Anthropic's use of books to train its AI models was lawful under fair use, but that conclusion came with a costly asterisk.

Alsup separately determined that Anthropic had obtained many of those books by pirating them from illegal shadow libraries rather than acquiring them legitimately. That distinction mattered enormously: the training itself was permissible, but the method of sourcing the material was not.

The practical consequence was steep. Anthropic was ordered to pay $1.5 billion to settle claims brought by a group of authors whose books ended up in its training pipeline.

Training purpose changes the analysis

A separate case drew a different line entirely. Judge Stephanos Bibas ruled that training an AI system on Reuters' content was not protected by fair use, because the resulting product was built to compete directly with Reuters in the same market.

Taken together, these decisions suggest courts are less concerned with the act of training itself and more focused on two separate questions: how the training data was obtained, and whether the resulting product substitutes for the original work in the marketplace. Lawful training on legitimately acquired material, as in Alsup's initial finding, is treated differently from training on pirated copies, and both are treated differently from building a direct competitor to the source, as in Bibas's ruling.

For AI companies, the emerging pattern is a set of narrower obligations rather than a single bright-line rule: source data legally, and be cautious about building products that directly displace the copyright holders whose work trained them. For publishers and authors, these rulings offer leverage, but not yet a settled legal framework, since Congress has not rewritten copyright law to address generative AI at all.

Original reporting and research used to synthesize this article.

  1. 1Is it legal to train AI models on copyrighted books? It’s complicatedtechcrunch.com
WK

WAKIB Editorial Team

This review was prepared and summarized by the WAKIB AI intelligence engine and vetted by our editorial board for accuracy and reliability.

Subscribe to Newsletter

Get a weekly summary of the most promising AI research and tools delivered to your inbox.

Telegram Channel

Join our active community on Telegram for real-time tracking of AI models and trends.

Join us on Telegram