TheVortiq
Inteligencia Artificial

Anthropic to Pay $1.5B for Using Pirated Books to Train Claude

Judge Approves Largest Copyright Settlement in History: 91% of Claimed Works

July 21, 2026 · 4 min read

wooden gavel and block on marble

TL;DR: A judge approved the $1.5 billion settlement between Anthropic and authors for using pirated books to train Claude. The ruling establishes that training with protected works may be fair use, but using pirated copies is not. It is the largest copyright settlement in history.

What Happened?

On Monday, federal judge Araceli Martínez-Olguín approved the $1.5 billion settlement between Anthropic and a group of authors and publishers who sued the AI startup for using copyrighted books to train its chatbot Claude. The settlement, described by plaintiffs' attorney Justin Nelson as "the largest copyright recovery in history," establishes that approximately 91% of the more than 482,000 books covered have been claimed, and rights holders will receive about $3,000 per work. This amount, while substantial, represents only a fraction of the $7 billion Anthropic has raised in funding, suggesting the cost is manageable for the company but sends a clear signal to the industry about the consequences of using unauthorized data.

Case Background

The lawsuit, initially filed in 2023, alleged that Anthropic used books obtained from shadow libraries such as Library Genesis and Sci-Hub to train Claude. These libraries host unauthorized copies of millions of protected works. The court determined that while training AI models with copyrighted data does not constitute infringement per se (under fair use), using pirated copies does violate the law. This nuance is crucial for the industry: it opens the door for other AI companies to use protected works as long as they obtain them legally, but closes the door to exploiting shadow libraries. This ruling adds to a series of judicial decisions defining the limits of fair use in the AI era. For example, in 2023, a judge dismissed similar lawsuits against OpenAI for lack of evidence, but this case sets a clearer precedent on the legality of data sources.

Significance of the Ruling

This settlement sets a significant precedent for the intersection of artificial intelligence and copyright. On one hand, it recognizes that training AI models can be considered fair use, giving breathing room to companies like OpenAI, Google, and Meta. On the other hand, it establishes that the source of data must be legal, forcing companies to audit their training datasets. The $1.5 billion payment, though high, is manageable for Anthropic, which has raised over $7 billion in funding. However, it sets a precedent for future lawsuits against other AI companies. Compared to the $9 million settlement Google paid in 2023 for using book data without permission, this amount is 166 times larger, reflecting the scale of the infringement and the growing legal pressure on AI companies.

Consequences for the Industry

The decision will have several repercussions:

  • Increased scrutiny of data sources: AI companies will need to demonstrate that their training data was obtained legally, which could increase data acquisition costs. According to industry estimates, cleaning and verifying datasets can increase training costs by up to 30%.
  • Promotion of licensing agreements: More companies are expected to sign agreements with publishers and authors, as OpenAI has already done with some media outlets (e.g., deals with Axel Springer and Associated Press). This could create a licensing market for training data, similar to that for music or images.
  • Impact on startups: Startups with fewer resources could be disadvantaged if they cannot access high-quality training data without paying licenses. This could consolidate the market in the hands of large companies with the ability to pay, reducing competition.
  • Debate on fair use: The ruling reinforces the fair use doctrine for AI training, but only if data is obtained legally. This could lead to further litigation over what constitutes a legal source, especially for data scraped from the internet where legality is ambiguous.

Reactions and Next Steps

The plaintiffs' attorney hailed the settlement as "a milestone for authors." Anthropic has not issued an official statement, but is expected to implement measures to ensure future datasets are legal. Judge Martínez-Olguín noted that the settlement provides "significant relief" to those affected. The case will continue to be monitored by the industry, as it could influence other similar lawsuits against AI companies. For example, there are pending lawsuits against OpenAI by authors such as George R.R. Martin and John Grisham, and against Meta for using books in training LLaMA. This settlement could serve as a model for resolving those cases, though the amount may vary depending on the scale of the infringement.

What Readers Should Know

This case demonstrates that artificial intelligence does not operate in a legal vacuum. Companies must respect copyright, even when innovating. For creators, it is a victory that recognizes the value of their work. For AI users, it is a guarantee that the models they use are built on legal foundations, though the debate over fair use continues. Additionally, this ruling underscores the importance of transparency in data sources: AI companies that do not audit their training datasets expose themselves to significant legal risks. Ultimately, this case could accelerate the creation of industry standards for data acquisition, benefiting both creators and AI developers.

Keep reading