NYT vs OpenAI: The Battle That Will Define Copyright in the AI Era
The New York newspaper has spent more than $20 million on a lawsuit that could change the rules of the game for language models.
July 28, 2026 · 4 min read

TL;DR: The New York Times sued OpenAI and Microsoft for copyright infringement, alleging they used its articles to train AI models. The case, which has already cost over $20 million, could change the rules of the game for training language models.
What happened?
In December 2023, The New York Times (NYT) filed a lawsuit against OpenAI and Microsoft for copyright infringement, alleging that the defendants used millions of the newspaper's articles without authorization to train their language models, such as GPT-4 and Copilot. According to the lawsuit, the chatbots reproduce verbatim excerpts from the articles, directly competing with the original content and harming the Times' business model. This case did not arise overnight: the NYT had attempted to negotiate a license with OpenAI for months, but talks broke down, leading to legal action. The lawsuit seeks statutory damages for massive infringement, as well as the destruction of any models that used the Times' data. OpenAI has responded that using public data to train AI is protected by the fair use doctrine, arguing that models do not store literal copies but learn statistical patterns. However, the NYT presented concrete examples of near-verbatim reproduction of paragraphs, strengthening its case.
Why is it important?
This is the first large-scale litigation between a traditional publisher and a leading generative AI company, and it could set a historic precedent. Until now, AI companies have operated in a legal gray zone, scraping data from the internet without compensating creators. The NYT case is especially significant because the newspaper has an archive spanning over 170 years and has heavily invested in quality journalism. If the Times wins, it could force OpenAI to pay retroactive licenses and establish a system of ongoing compensation for publishers, similar to what happened with digital music (like Spotify's deals with record labels). If it loses, it would open the door for AI companies to continue using protected content without payment, threatening the viability of professional journalism. Sulzberger himself has stated he is willing to go to the Supreme Court to defend copyright. Additionally, this case joins other similar lawsuits, such as Getty Images v. Stability AI for using its photos without a license, and several authors (including George R.R. Martin and John Grisham) against OpenAI for copyright infringement in books.
Potential consequences
- Economic impact: OpenAI could face billions of dollars in damages. The NYT alone has spent over $20 million in legal fees so far, and the legal battle could last years. Moreover, if licenses are required, the cost of training language models could skyrocket, affecting startups and AI innovation. On the other hand, publishers could gain a new revenue stream through licensing deals, as OpenAI has already done with Associated Press, Axel Springer, and Le Monde, paying undisclosed sums estimated in the millions of dollars annually.
- Legal precedent: The ruling could define the limits of fair use in the AI era. So far, courts have been favorable to fair use in cases like Google Books (which allowed scanning books for search), but reproducing creative content by generative models is new territory. The NYT case could establish that training models with protected data without permission is not fair use, especially if the model can generate content that directly competes with the original.
- Business model: If the Times wins, we are likely to see a wave of similar lawsuits from other publishers, such as The Wall Street Journal, The Guardian, or Reuters. This could lead to the creation of a data licensing market for AI, where publishers charge for the use of their content. However, it could also incentivize AI companies to seek synthetic or public domain data, reducing training quality.
What should readers know?
The case is currently in the discovery phase, where both parties are exchanging documents and testimonies. The trial is expected to begin in late 2024 or early 2025, but could be prolonged for years if there are appeals. The NYT has requested that all OpenAI models trained on its data be destroyed, which would be an unprecedented measure. OpenAI, for its part, has argued that the Times is exaggerating and that its models do not store literal copies. Additionally, they have pointed out that the newspaper used prompts specifically designed to induce the models to reproduce text, which does not reflect normal use. Other key players, such as Microsoft, have backed OpenAI, while copyright groups like the Authors Guild have supported the Times. This case is part of a broader debate about the ethics and legality of web scraping for AI training. Unlike music or video, where licenses are well-established, text content still lacks a clear framework. The resolution of this case could radically change how AI companies access training data, affecting not only publishers but also researchers, developers, and end users. For now, readers should watch for court decisions, which could come sooner than expected if there is a summary judgment.