Skip to content

Training models on copyrighted books: US courts draw line between reading and piracy amid fair use uncertainty

Share
Training models on copyrighted books: US courts draw line between reading and piracy amid fair use uncertainty

Listen to this article

Read by Anchor

Large language models that power AI tools such as ChatGPT, Gemini and Claude rely on massive databases containing hundreds of millions of books, published articles and academic papers available on the internet. Although most authors feel that their works have been used without permission to build tools that could undermine their professional livelihoods, the legal reality surrounding the legitimacy of training algorithms on copyrighted material appears highly complex and tangled within U.S. courts.

U.S. courts have already begun drawing a line between consuming content and pirating it.In a historic judicial ruling, Judge William Alsup ordered the company Anthropic to pay a settlement of $1.5 billion to a group of writers. While the settlement appeared to be a symbolic victory for the authors, Alsup noted in the opinion that training AI models is itself a lawful activity, pointing out that the models train on texts not to copy them or replace them, but to generate different content, much like an aspiring writer who studies literature to hone their craft. The financial penalty resulted from the company obtaining the works through illegal digital libraries.

Intellectual-property lawyer Kathy Gillis believes this judicial approach represents a real gain for AI companies, especially given that courts rely on the 1976 Copyright Act, which has not undergone a major amendment in half a century. Gillis explains that the law penalizes the act of physical copying but does not prohibit reading a work, testing it or its cognitive consumption, which makes training resemble human reading for companies aiming to record annual revenues of roughly $200 billion by 2028.

The fair-use standard and the limits of commercial competition serve as a compass for current disputes.Attorney Jason Henderson notes that courts tend to permit training when the use is transformative and does not target direct competition with the original product. In this context, the Thomson Reuters case against Ross Intelligence stands out, where Judge Stephanos Pipas ruled that fair use does not apply to copying Reuters content to build a legal platform that competes with it directly.

The legal debate also extends to the models’ outputs, as earlier rulings such as the Thaler v. Perlmutter case have held that works generated entirely by AI are denied copyright protection, opening an ongoing discussion about the line between technical assistance and human authorship. With pending cases continuing, the legal landscape remains in a formative stage, compelling sector companies to track every precedent to avoid forthcoming regulatory risks.

Don't miss the next story

Subscribe for updates