The New York Times and The Daily News claim OpenAI hid training data and chat logs that could reveal copyrighted journalism, prompting a fresh motion for sanctions in their two‑year legal battle.
The New York Times and The Daily News have jointly accused OpenAI of deliberately concealing critical evidence in the ongoing copyright lawsuit. According to the publishers, the AI giant maintained a massive internal database of roughly 78 million de‑identified ChatGPT conversations and employed specialized tools to locate copyrighted journalism, all while claiming it could not search its own training corpus. This allegation intensifies a two‑year fight over whether OpenAI’s generative models unlawfully trained on the newspapers’ content and reproduced it in user‑facing outputs.
Plaintiffs’ Demands and OpenAI’s Defense
The media outlets seek concrete proof that their copyrighted articles were incorporated into OpenAI’s training data and that ChatGPT has repeatedly reproduced those pieces. OpenAI, for its part, has consistently argued that it lacks the technical ability to search its vast training set and that extracting, processing, and de‑identifying billions of chat logs would be both burdensome and a privacy risk for users.
Undisclosed Findings and the “Project Giraffe” Reveal
During an April court‑ordered deposition, OpenAI data‑privacy engineer Vinnie Monaco allegedly admitted that the company had already conducted internal searches for copyrighted journalism. Monaco’s testimony further indicated that, even before the lawsuit was filed, OpenAI had assembled a database of about 78 million de‑identified ChatGPT conversations to gauge the extent of potential infringement. In addition, the firm reportedly deployed a “Bloom” filter within a suite of tools dubbed “Project Giraffe,” designed to detect and log instances of content regurgitation in model outputs.
Questions Over the Reliability of Court‑Submitted Samples
Originally, the plaintiffs demanded a sample of 120 million chat logs. OpenAI negotiated that number down to 20 million, which it finally submitted in December. However, the court described the sample as “unusable” because of extensive redactions. The publishers also allege that OpenAI deleted billions of ChatGPT outputs after the lawsuit was filed—directly contravening a preservation order—and substituted millions of logs in the sample it provided, effectively making the evidence unattainable.
Legal Motions and Future Implications
Now, the NYT and The Daily News are urging the judge to sanction OpenAI for withholding evidence and obstructing discovery. They request that the 20 million‑log sample be excluded as unreliable, that the court accept that the logs would have shown substantial regurgitation of the plaintiffs’ content, and that OpenAI be ordered to pay the plaintiffs’ legal costs. OpenAI spokesperson Drew Pusateri denied the accusations, accusing the newspapers of attempting to pry into private user conversations as their case weakens. “We will continue to defend our users’ privacy and the long‑standing principles of fair use,” he stated.
This dispute spotlights the growing tension between AI developers, copyright holders, and privacy advocates. How courts balance the need for transparency in AI training with the protection of user data will shape the regulatory landscape for generative AI for years to come.