Anthropic pays $1.5B for using pirated books in AI training
The largest copyright settlement in the U.S. sets a historic precedent for the AI industry
July 24, 2026 · 4 min read

TL;DR: Anthropic will pay $1.5 billion to authors for training its Claude models on pirated books, in the largest copyright settlement in the U.S. This sets a precedent that will force the entire industry to license data legally.
What happened?
On July 20, 2026, Judge Araceli Martínez-Olguín of the San Francisco District Court gave final approval to the settlement under which Anthropic will pay $1.5 billion to a group of authors who accused the startup of using pirated books to train its Claude language models. This is the largest copyright infringement settlement in U.S. history, far surpassing the $500 million Google paid for its digital library. The case, known as Kadrey et al. v. Anthropic, was filed in 2024 by a group of authors including Sarah Silverman, Christopher Golden, and Richard Kadrey, who alleged that Anthropic used datasets such as The Pile, which included copyrighted works obtained from pirated sources like Bibliotik, Library Genesis, and Z-Library, to train its Claude models. According to court documents, more than 200,000 copyrighted books were identified in the training data. The settlement, approved without opposition, includes a $1.5 billion payment to a fund for affected authors, as well as Anthropic's commitment to implement a license verification system for future datasets.
Why is this important?
This case not only affects Anthropic but also sets a legal precedent for the entire industry. Until now, many AI companies argued that using public data for training constituted 'fair use.' However, the court determined that using copyrighted works obtained from pirated sources like Bibliotik, Library Genesis, or Z-Library is not protected by that doctrine. The decision is based on the U.S. Supreme Court's 2023 ruling in Andy Warhol Foundation v. Goldsmith, which restricted the scope of fair use for transformative works. Additionally, the court highlighted that Anthropic knew the data came from illegal sources, which aggravated its liability. This case is the first in a series of similar lawsuits: OpenAI faces a class action from authors represented by the Authors Guild, while Meta and Stability AI are also being sued for using protected data. The total value of potential settlements could exceed $10 billion, according to Bloomberg Law analysts.
"This settlement sends a clear message: if AI companies want to use protected material, they must pay for it," the judge stated during the approval hearing.
Consequences for the industry
The impact extends across multiple fronts:
- Compliance costs: AI startups will need to invest in license verification and data provenance systems, increasing operational costs. Companies like Hugging Face have already announced tools to audit datasets. Compliance costs for a mid-sized startup are estimated to increase by 20% to 30%.
- Business model restructuring: Companies like OpenAI, Meta, and Google are already renegotiating agreements with publishers and content platforms. OpenAI has signed deals with Axel Springer and Associated Press, while Meta is in talks with academic publishers. Google has expanded its agreement with Reddit to include training data.
- Greater transparency: Regulators are expected to require companies to disclose their training data sources. The European Union is already considering a directive similar to the Data Act, which would require transparency in datasets used by AI. In the U.S., the FTC has launched an investigation into AI companies' data collection practices.
- Impact on open models: Projects like Meta's LLaMA and Stability AI's Stable Diffusion could be affected if they cannot demonstrate the legality of their data. Meta has already had to withdraw some open-source models after infringement allegations.
What should readers know?
For users of AI tools, this case does not imply immediate changes in access or functionality. However, in the long term, it could lead to more expensive subscriptions if companies pass on licensing costs. For example, OpenAI has already increased ChatGPT Plus prices by 20% in some markets. It could also slow innovation in open models if independent developers cannot afford copyright fees. On the other hand, authors and content creators see this ruling as a historic victory that recognizes the value of their work in the digital age. Organizations like the Authors Guild have already announced they will seek similar settlements with other AI companies. For investors, the case signals growing regulatory risk: AI startups could face significant liabilities if they do not secure their data. Companies like Anthropic have already seen their valuation drop by 15% following the settlement, according to PitchBook.
The future of AI and copyright
This case is just the first of many. Similar lawsuits are ongoing against OpenAI, Meta, and Stability AI. Judge Martínez-Olguín's decision could influence those proceedings, though each case has its own specifics. What is clear is that the era of unrestricted 'fair use' for AI has ended. Companies will have to adapt to a new ecosystem where training data is negotiated like any other resource. Markets for training data are expected to emerge, similar to image libraries, where creators can license their works for AI. Companies like Shutterstock and Getty Images already offer licensed datasets. Additionally, insurers are developing specific policies to cover copyright risks in AI. In summary, this case marks a turning point in the relationship between artificial intelligence and intellectual property, with implications that will be felt for years.