News outlets are fighting OpenAI in court to stop the tech giant from deleting data they claim is crucial evidence of copyright infringement. Publishers argue the data proves OpenAI plagiarized content for its AI models, while OpenAI cites burden and user privacy concerns.

In a Manhattan federal courtroom, a high-stakes legal drama is unfolding.
It pits the venerable institutions of the Fourth Estate against the technological behemoth, OpenAI.
At the heart of the current skirmish is not just the colossal sum of alleged damages, nor the fundamental question of copyright in the age of artificial intelligence.
Instead, it is a far more immediate and visceral battle over data itself.
Specifically, the dispute concerns whether OpenAI can continue to delete information that news outlets claim is crucial evidence of plagiarism.
Lawyers representing titans like The New York Times and The Daily News, alongside a consortium of other prominent publications, have implored Magistrate Judge Ona Wang to reject OpenAI’s fervent plea.
They ask her to reject the request to resume the mass deletion of data.
This data, they argue, holds the key to proving that ChatGPT’s parent company has systematically pilfered copyrighted journalistic works.
They claim it circumvented paywalls, and then regurgitated this content, often distorted or misrepresented, for its own commercial gain.
The current judicial tug-of-war began last month.
Judge Wang, responding to accusations that OpenAI was intentionally purging “enormous swaths of data,” ordered the tech giant to preserve its output logs.
She also ordered the preservation of any related information slated for deletion.
The news outlets contend that these deletions were a deliberate attempt to obscure the paper trail of their intellectual property being ingested and repurposed by OpenAI’s large language models.
OpenAI, however, has countered.
It asserts that maintaining such a vast reservoir of data would constitute a “massive burden.”
Rather curiously, it also claims this would infringe upon the privacy of its users.
This argument has been met with a collective raised eyebrow from the plaintiffs.
They highlight the apparent contradiction between OpenAI’s public assurances to users.
Those assurances stated that data will be retained if legally required.
This contrasts with its current stance in court.
More pointedly, the news organizations underscore that OpenAI has yet to dispute the relevance of the data it seeks to expunge.
“What it does not dispute is that the output log data is relevant to the News Cases,” lawyers for the outlets wrote.
They added a stinging observation: “Nor can it dispute that, as a highly sophisticated technology company that is currently valued at more than $300 billion, it has both the means and ability to preserve this concededly relevant data.”
The implication is clear.
A company valued in the hundreds of billions of dollars, a pioneer in data processing and storage, is suddenly crying foul over the “burden” of retaining relevant information for a lawsuit.
This perceived evasiveness, the news outlets contend, is part of a pattern.
They accuse OpenAI of employing “every trick in the book” to skirt accountability.
This includes the alleged installation of filters designed to make it harder to elicit answers containing copyrighted journalistic works.
“OpenAI’s preferred course of action to ‘protect its users’ data and privacy’ — immediately resuming mass deletions — will also, coincidentally, allow it to continue to destroy data that would show its liability for copyright infringement,” their legal brief states.
This underscores a deep skepticism about the company’s motives.
Judge Wang, anticipating privacy concerns, had already meticulously outlined her May 13 order.
She stated that the preserved information would be segregated.
It would not be provided “wholesale” to anyone, nor stored “forever.”
It would be used solely to address the specific concerns raised in the lawsuit.
Should the judge even entertain OpenAI’s objection, the newspapers have urged her to allow them to analyze different populations of data.
They want to present their findings to the court, ensuring transparency in the face of alleged obfuscation.
The broader lawsuit paints a stark picture of modern intellectual property theft.
It alleges that OpenAI has illicitly harvested millions of news stories, the very lifeblood of journalistic enterprise, to train its generative AI products.
These products, the suit claims, then “vomit out” these stories, or versions thereof, to users.
This process, the newspapers argue, not only bypasses their paywalls.
It often results in journalists’ pirated reporting being misstated or misrepresented.
Thereby, it misinforms ChatGPT users and erodes public trust.
The disparity in effort and investment is a central theme.
While news publishers collectively spend billions dispatching “real people to real places to report on real events in the real world,” the lawsuit contends that tech firms like OpenAI are “purloining” this painstakingly gathered reporting without compensation.
The objective, they argue, is to create products that directly compete.
These products provide “news and information plagiarized and stolen.”
OpenAI, for its part, has consistently cloaked its operations in the broad mantle of “fair use.”
This is a legal doctrine traditionally applied to transformative works like criticism, commentary, news reporting, teaching, and research.
However, lawyers for the newspapers vehemently dispute this applicability.
They argue that the fair use test requires a copyrighted work to be transformed into something genuinely new.
Critically, they state that this new work cannot compete with the original in the same marketplace.
By generating news summaries or articles based on their copyrighted content, OpenAI’s products, they argue, clearly compete with and undermine the original source.
Adding weight to the news outlets’ claims, the judge has already rejected OpenAI’s prior assertion.
That assertion was that the newspapers haven’t produced “a shred of evidence” that people are using ChatGPT or OpenAI’s API products to get news instead of paying for it.
The plaintiffs recently pointed out that OpenAI’s own engineers have all but admitted as much.
They acknowledged that while chatbots weren’t designed to slip past paywalls, that doesn’t mean they couldn’t.
They also cited a separate suit involving Google.
There, an OpenAI engineer conceded that local news was a “pretty common quer[y]” among ChatGPT users.
This was a revealing admission for a company claiming no competitive intent.
The legal battle commenced with The New York Times filing its suit in Manhattan Federal Court in December 2023.
The Daily News, alongside MediaNews Group and Tribune Publishing affiliates, joined the fray in April 2024.
These affiliates include The Mercury News, The Denver Post, The Orange County Register, the St. Paul Pioneer Press, the Chicago Tribune, the Orlando Sentinel, and the South Florida Sun Sentinel.
As this pivotal legal fight continues to unfold, OpenAI’s lawyers have remained notably silent.
They have declined requests for comment from The News.
This leaves the courtroom to speak for itself in this defining moment for intellectual property in the AI age.