AI Training on Copyrighted Books Remains Legally Unsettled
The legality of training AI models on copyrighted books remains unresolved despite a series of high-profile court rulings, according to a report by TechCrunch.
TechCrunch spoke with intellectual property attorneys Cathy Gellis and Jason Henderson about the state of AI copyright litigation. Last year, Judge William Alsup ordered Anthropic to pay a $1.5 billion settlement to a group of writers, but ruled that the underlying AI training itself was lawful. Anthropic was penalized specifically for pirating books from illegal shadow libraries, not for the training process. Alsup compared an AI model ingesting text to a writer studying literature.
Gellis told TechCrunch the ruling is largely favorable to AI companies, noting that a $1.5 billion penalty is minor relative to industry revenue projections. She said the decision treated AI training as more analogous to reading a work than copying it, since copyright law centers on copying rather than consumption.
Henderson pointed to a separate case, Thomson Reuters v. Ross Intelligence, in which Judge Stephanos Bibas ruled against Ross for training on Reuters content to build a directly competing legal research platform, finding the use was not transformative. Henderson said courts have tended to favor AI developers when the resulting product does not directly compete with the original copyrighted material, and to rule against them when it does.
TechCrunch also noted the related but distinct issue raised in Thaler v. Perlmutter, in which a court found that fully AI-generated works are not copyrightable, raising further questions about how to assess AI-assisted content.
According to TechCrunch, most major AI companies remain in ongoing litigation over these issues, and attorneys say a definitive legal standard is unlikely to emerge soon, as early rulings continue to be tested and potentially overturned in later cases.
Based on reporting by techcrunch.com.
