Is Training AI on Copyrighted Books Legal? Inside the Complex Court Battles
Artificial intelligence models powering systems like ChatGPT, Gemini, and Claude rely on massive datasets of published works, including hundreds of millions of books, academic papers, and articles. While many authors raise concerns that their works are being utilized without explicit consent, the legal boundaries governing AI training remain intricate and evolving across courtroom decisions.
What Happened
Legal battles between copyright holders and AI companies have begun establishing early precedents. In a prominent case, Judge William Alsup ordered AI firm Anthropic to pay a $1.5 billion copyright settlement to a group of authors. However, the ruling established that the act of training an AI model itself was lawful. The court penalised Anthropic specifically for acquiring the books through illicit online shadow libraries rather than for the underlying model training.
Judge Alsup compared an AI model ingesting text to an aspiring writer studying literature to create something new rather than replicating existing texts. Legal experts note that this interpretation positions model training closer to reading or consuming a work rather than infringing on copyright protections, which traditionally regulate unauthorized copying.
Key Highlights
- Judge William Alsup ruled that Anthropic’s AI training was lawful, penalising the company $1.5 billion strictly for acquiring books from illegal shadow libraries.
- United States copyright law has not undergone a major update since 1976, forcing modern courts to apply decades-old frameworks to complex generative AI technologies.
- Fair use assessments often turn on whether the use is transformative and whether the AI system directly competes with the original copyrighted work in the marketplace.
- In Thomson Reuters v. Ross Intelligence, Judge Stephanos Bibas found that training on proprietary content to build a directly competing legal research platform was not fair use.
- In Thaler v. Perlmutter, the court established that completely AI-generated works cannot receive copyright protection.
Why This Matters
The distinction between consuming material and copying it is critical for the future of generative technology. Intellectual property attorney Cathy Gellis explained that copyright law targets copying rather than reading or experiencing a work. Because Anthropic projects approximately $200 billion in annual revenue by 2028, legal settlements focused narrowly on piracy rather than the core training process leave the AI training model viable.
Market competition plays a defining role in judicial reasoning. According to Jason Henderson, founder of the IP and Media Practice at JWL International, courts examine whether the training purpose directly challenges the original material. While courts rejected fair use when Ross Intelligence copied Thomson Reuters content to launch a competing platform, authors claiming that chatbots supplant their written books have not yet prevailed with that argument.
What to Watch Next
Numerous major AI developers remain involved in pending litigation over data usage, fair use, and synthetic content generation. Legal observers note that current decisions represent opening rulings that could be shaped or overturned as higher courts and subsequent litigation phases address the intersection of artificial intelligence and intellectual property.
Frequently Asked Questions
Why was Anthropic fined if AI training was considered lawful?
Anthropic was penalized for sourcing books from illegal online shadow libraries. The court treated the training process itself as lawful transformative learning, separating the source piracy from the mechanics of AI ingestion.
How does copyright law handle fair use in AI training?
Fair use assessments depend on the purpose and character of the work, the amount used, and the impact on the existing market. Courts tend to disallow training when it serves to build a tool that directly competes with the copyrighted source material.
Can content generated entirely by AI be copyrighted?
Under the Thaler v. Perlmutter ruling, works that are 100 percent generated by artificial intelligence are not eligible for copyright protection.
Source: TechCrunch.
