
Newly unsealed court filings in The New York Times’ copyright lawsuit against OpenAI and Microsoft have revealed internal comments from executives and employees that could add fresh pressure to the companies’ defenses over how copyrighted news content was used to develop artificial intelligence systems. The documents include a January 2023 internal memo in which Microsoft director of Applied Science Brent Hecht described the scale of AI scraping as “an astonishing theft of unprecedented proportions” and “the largest theft of labor in human history,” according to material cited in the Times’ latest court filing.
Other newly disclosed material includes remarks from OpenAI executives about the threat AI products could pose to publishers, internal discussions about accessing paywalled material and data showing that Microsoft’s Copilot could sharply reduce traffic to news websites.
The revelations do not amount to a court finding that Microsoft or OpenAI committed copyright infringement. The underlying legal dispute remains unresolved, and the companies continue to defend their use of copyrighted material as lawful.
What the new filings reveal about AI scraping
The Times lawsuit alleges that OpenAI and Microsoft copied millions of news articles without permission and used them in developing and operating generative AI systems.
The newly unsealed material, much of which is presented through arguments made by the Times and other publishers, provides a more detailed look at how some of that content was collected.
According to the filings, OpenAI used large web-scraped datasets and, in some instances, accessed content that was behind the Times’ paywall. The documents also allege that copyright notices were removed from some training data before it was used to develop models.
Those allegations are central to the publishers’ argument that the companies’ use of the material was not simply an abstract technical exercise, but part of a commercial system that could compete directly with the original publications.
Microsoft’s own data showed a steep traffic decline
One of the most consequential pieces of evidence concerns Microsoft Copilot.
According to an internal Microsoft presentation from January 2024 cited in the filings, Copilot’s “answer engine” reduced click-through rates to The New York Times website by as much as 93% compared with traditional Bing search.
The Microsoft document described the situation as a “doom loop,” warning that reduced traffic could ultimately damage both the publishers supplying content and the AI systems that depend on that content.
This issue goes to the heart of copyright law’s market-impact analysis.
The publishers argue that if AI systems provide information directly to users instead of directing them to the original article, those systems can reduce the economic value of the underlying journalism.
OpenAI, however, has argued that generative AI training is transformative and does not simply function as a substitute for the copyrighted works used to train the models.
Satya Nadella’s testimony adds another layer
Microsoft CEO Satya Nadella also made significant comments during a deposition earlier this year.
According to the newly cited court material, Nadella said that “anything that is paywalled” should be licensed by organizations using it for “grounding or training.”
He also testified that had he known OpenAI was scraping and training on paywalled material, he would have invoked Microsoft’s contractual rights to require OpenAI to retrain its models..
The testimony does not necessarily establish Microsoft’s corporate legal position on every form of AI training. Microsoft continues to argue that certain uses of copyrighted material can qualify as fair use.
But Nadella’s remarks could become important for the publishers as they argue that even within the companies, there was recognition that access to paywalled content raised distinct licensing concerns.
OpenAI executive called AI products an “existential threat”
The filings also contain comments attributed to OpenAI’s Head of ChatGPT, Nick Turley.
According to the Times’ court submission, Turley described products such as ChatGPT as an “existential threat” to publishers and said they were “largely substitutive,” with that substitution likely to increase as the technology improved.
The language is significant because one of the key fair-use questions is whether a new use harms or substitutes for the market for the original work.
OpenAI’s public legal position is that AI models do not simply reproduce copyrighted works and that training is a transformative analytical use that can coexist with, rather than replace, original content.
The publishers are using the newly disclosed statements to argue that internal discussions recognized the possibility of direct economic substitution.
How much news content was copied?
The filings also highlight the scale of the data involved.
They state that OpenAI’s mid-training datasets contained more than 91,692 copies of works from The New York Times, Daily News and the Center for Investigative Reporting.
Another Common Crawl-derived dataset allegedly contained more than 2 million documents from nytimes.com alone.
The documents also describe Project Mango, a dataset assembled through collaboration between Microsoft and OpenAI that allegedly included copies of at least 160,903 unique works from news publishers.
These numbers illustrate the central issue in the litigation: AI training can involve copying content on a scale far larger than the use of individual articles by a conventional reader.
The legal question is whether that type of copying, when used to train AI systems, is protected by copyright law.
The paywall question
Perhaps the most controversial allegation involves efforts to access articles that were not freely available online.
The filings describe internal OpenAI discussions about a technique for getting around The New York Times’ paywall. According to the material cited by the publishers, OpenAI researcher Nick Ryder informed OpenAI president Greg Brockman about a method to bypass the restriction, to which Brockman responded positively.
The publishers argue that such evidence undermines attempts to characterize all of the training data as material legitimately available for computational analysis.
The companies’ legal defenses, however, involve broader questions about fair use and the technical processes used to train AI systems, not simply whether individual employees accessed specific articles.
Microsoft and OpenAI shared training data
The documents also shed light on the close relationship between the two companies.
According to the filing, OpenAI provided Microsoft with the entire GPT-3 training dataset so Microsoft could evaluate how to incorporate OpenAI models into its commercial products.
Microsoft also supplied training data to OpenAI through projects referred to as Project Taxi and Project Mango.
The disclosures add detail to a partnership that has become central to the development and commercialization of generative AI, while also illustrating how training data can move between companies involved in building and deploying the models.
Why this matters beyond The New York Times
The case is bigger than a dispute between one newspaper and two technology companies.
Publishers around the world are grappling with a basic economic question: if AI systems can answer a user’s question directly using information derived from journalism, what happens to the audience, advertising and subscription revenue that traditionally support that journalism?
The Times and other publishers contend that AI companies should obtain permission or licenses to use their work. OpenAI maintains that AI training can qualify as fair use and says its systems have transformative purposes that differ from simply republishing source material.
Those competing arguments are now being tested in courts across the United States.
What does this mean for the fair-use fight?
US copyright law generally considers four factors when determining whether an unlicensed use may qualify as fair use, including the purpose of the use, the nature of the copyrighted work, how much was used and the effect on the potential market for the original.
The newly disclosed documents are particularly relevant to the first and fourth factors because they contain internal discussions about commercial AI products, the use of copyrighted news and possible effects on publishers’ traffic and business models.
But no single internal email or presentation decides the legal question.
Courts must evaluate the evidence as a whole, and the final outcome will depend on how judges interpret the technology and the facts under US copyright law.
A major test for the AI industry
The newly unsealed material provides an unusually detailed glimpse into the tensions created by the AI industry’s dependence on vast quantities of online information.
For Microsoft and OpenAI, the lawsuit is part of a broader legal battle over whether existing copyright rules permit the use of protected material to develop increasingly powerful AI models.
For publishers, the stakes are similarly large. Their argument is not simply about ownership of individual articles, but about whether a business model built on decades of reporting can survive when AI systems can absorb and deliver that information without necessarily sending users back to the source.
The latest filings do not settle that debate. They do, however, reveal just how intensely companies building AI have been thinking about the same economic problem confronting the publishers whose work helped create the information ecosystem on which those systems depend.