Artists vs. AI: The Legal Battle Redefining Copyright
Writers, illustrators, and photographers take Google, Meta, and Anthropic to court for training models with their works without permission or compensation.
August 1, 2026 · 5 min read
TL;DR: Artists have sued Google, Meta, and Anthropic for training AI models with their protected works without authorization. The cases could redefine fair use and force companies to pay licenses.
What Happened?
In July 2025, the release of a dataset by The Atlantic allowed anyone to search which copyrighted works had been used to train artificial intelligence models. Dozens of artists—writers, illustrators, musicians, and photographers—discovered that their creations appeared in datasets such as Common Crawl, LAION-5B, or pirated books, without their consent or compensation. This triggered a wave of class-action lawsuits against tech giants like Google, Meta, Anthropic, OpenAI, and Stability AI.
Kirk Wallace Johnson, author of The Feather Thief and The Fishermen and the Dragon, told The Verge: “I felt a mix of anger at the audacity of the theft, concern about what this means for writers, and a healthy thirst for revenge against these corporations that have become galactically rich” using his work. Johnson spent “five or six years researching, writing, and researching” his books, only to find them pirated and feeding chatbots.
This phenomenon is not new: in 2023, authors like Sarah Silverman and George R.R. Martin had already sued OpenAI for copyright infringement. However, the current magnitude is greater: according to data from the Authors Guild, more than 15,000 writers have joined class-action lawsuits, and organizations like the Illustrators' Partnership estimate that 70% of professional illustrators have seen their works in training datasets without permission.
Why It Matters
These cases are not isolated incidents. They represent a direct challenge to the business model of generative AI, which relies on vast amounts of data to train its systems. According to a 2024 report from Stanford HAI, 58% of datasets used to train language models contain copyrighted material. If courts rule in favor of artists, companies could be forced to pay retroactive licenses, remove datasets, or redesign their models from scratch, costing billions of dollars.
Moreover, the outcome could define the scope of fair use in the context of AI training. Until now, companies have relied on this U.S. legal doctrine, which allows limited use of protected material without permission for purposes such as research or criticism. However, plaintiffs argue that the massive commercial use of their works to create products that directly compete with them cannot be considered “fair use.” A key precedent is the Google Books case (2015), where the Second Circuit Court of Appeals considered the digitization of books for search to be fair use. But generative AI goes further: it not only indexes but generates new content that can displace original creators.
The economic impact is massive. According to a Goldman Sachs study (2024), generative AI could automate up to 25% of content creation tasks, threatening the income of writers, musicians, and visual artists. The International Federation of the Phonographic Industry (IFPI) reported that AI-generated music already accounts for 8% of streams on platforms like Spotify, without original artists receiving compensation.
Potential Consequences
- For AI companies: Risk of million-dollar damages and changes in their data collection practices. Some have already begun signing licensing agreements with publishers and stock image banks. For example, OpenAI paid Shutterstock for access to its images and signed a deal with Axel Springer for news. Meta, for its part, has announced it will remove problematic datasets like Books3, but has not yet compensated authors. According to The Verge, Google has allocated $100 million to a copyright fund, though critics consider it insufficient.
- For artists: Possibility of receiving compensation for past use of their works and establishing a framework for future licenses. The Authors Guild has proposed a collective licensing system similar to music rights management societies. However, the legal process is slow: class actions can take years to resolve, and meanwhile, many artists see their incomes reduced. For example, illustrator Sarah Andersen, a plaintiff in a case against Stability AI, reported a 40% drop in her sales since image generators became popular.
- For the market: It could slow the development of new models if access to data is restricted, or conversely, incentivize the creation of ethical and transparent datasets. Companies like Adobe have launched models trained only on licensed content (Firefly), while startups like Fairly Trained offer data ethics certifications. If the lawsuits succeed, we are likely to see an increase in the cost of training data, which could consolidate the market in the hands of giants that already have agreements, such as Microsoft (which invested in OpenAI) and Google (which owns its own data ecosystem).
What Readers Should Know
Not all lawsuits have been successful. Some have been dismissed for lack of concrete evidence of harm or because the transformative use of the AI model was considered protected. For example, in 2024, a judge dismissed a lawsuit against Google over the use of books in its PaLM model, arguing that specific economic harm was not demonstrated. However, cases like Getty Images v. Stability AI (over the use of its photos in Stable Diffusion) have advanced: in January 2025, a British court allowed the case to proceed, indicating that courts are taking the issue seriously.
For creators, it is crucial to document their intellectual property and monitor platforms like Have I Been Trained? (developed by digital artist Laion) to know if their works have been included in datasets. They can also opt for protection tools like Glaze or Nightshade, which alter images to confuse AI models. For companies, the lesson is clear: opacity in training sources is no longer a viable option. The European Commission has already included transparency requirements for datasets in its AI Act, and in the United States, the Copyright Office is considering new regulations.
“It's not about slowing innovation, but about ensuring innovation is not built on the exploitation of others' work,” summarizes copyright lawyer Sarah Jeong, former editorial board member of the New York Times.
In short, these lawsuits are just the beginning of a long process that will define the balance between technological advancement and the protection of creators' rights in the digital age. Generative AI promises to transform entire industries, but the fundamental question remains: at what cost and for whom?