NE Times
Opinion

The Real Story in NYT v OpenAI Is No Longer Copyright. It Is Candour.

Publishers now accuse OpenAI of misleading the court about what it could search. Whoever wins, the discovery fight will define trust between AI and journalism.

The NE Times Editorial Board

Commentary & Analysis ·

3 min read
Illustration of scales weighing a glowing AI network against a stack of newspapers before a wall of court files

Verified key facts

  • On 9 July 2026, the New York Times, the New York Daily News and other publishers asked a Manhattan federal court to sanction OpenAI in their copyright case, per the Washington Post and TechCrunch.
  • The publishers allege OpenAI falsely told the court it could not search its systems for their copyrighted material, while having run such internal searches even before the Times filed suit, per TechCrunch.
  • An April deposition of OpenAI engineer Vinnie Monaco allegedly revealed OpenAI had amassed a database of about 78 million de-identified ChatGPT conversations used internally to assess infringement, per TechCrunch.
  • Requested sanctions include excluding a 20-million chat-log sample, adverse factual findings on regurgitation, and legal fees, per Bloomberg Law and TechCrunch.
  • OpenAI spokesperson Drew Pusateri called the allegations 'blatantly false' and accused the Times of seeking to invade the privacy of uninvolved users, per the Washington Post.

A copyright case becomes a candour case

The most consequential media lawsuit of the decade changed character last week. On 9 July, a group of publishers led by the New York Times asked a Manhattan federal judge to sanction OpenAI, as the Washington Post reported. The claim is stark: that OpenAI told the court it could not search its systems for evidence of copyrighted journalism in its training data, while it had already done exactly that.

According to TechCrunch, an April deposition of OpenAI engineer Vinnie Monaco revealed internal searches of the training corpus that began before the Times even filed suit. The publishers also say OpenAI built a database of roughly 78 million de-identified ChatGPT conversations to measure, internally, how much it was reproducing others' work. OpenAI denies the allegations. Spokesperson Drew Pusateri called them 'blatantly false' and accused the Times of trying to invade users' privacy as its case weakens.

Why discovery conduct matters more than the verdict

Courts have not yet ruled on the merits of the copyright claims. But the sanctions fight may matter more, because it tests something the AI industry constantly asks of the public: trust us. AI companies hold all the evidence about what their models ingested and what they emit. Publishers, regulators and users can only see what those companies choose to disclose. Litigation discovery is one of the few mechanisms that forces the curtain open.

If a federal court finds that the industry's most prominent firm misrepresented its own capabilities to a judge, the damage will not stay inside this docket. Every future negotiation over licensing, every regulatory filing, every safety attestation gets read differently. Candour under oath is the cheapest trust infrastructure that exists. Squandering it is expensive for everyone.

OpenAI's defence deserves a fair hearing

There is a serious version of OpenAI's position. Searching a training corpus is not one thing; it spans many systems, formats and eras, and a company can honestly say one kind of search is infeasible while another kind exists. The privacy objection is not frivolous either. The publishers' demands do implicate logs of millions of users with no stake in this dispute, and 'de-identified' data is notoriously re-identifiable.

The judge, not the press, should weigh those defences. But note what the defence concedes: that the company alone decides which of its internal capabilities the outside world learns about. That asymmetry, not any single misstatement, is the structural problem this case has exposed.

What this means for the economics of news

The requested sanctions, per Bloomberg Law and TechCrunch, include excluding a 20-million conversation sample and adverse findings that chat logs would have shown substantial regurgitation of the plaintiffs' content. If granted, they would strengthen the publishers' core economic argument: that AI systems substitute for the journalism they were trained on.

That argument is the whole game for the news business. Three outcomes are plausible from here.

  • A publisher win or strong settlement, which would set a de facto price on news in training data and accelerate licensing markets.
  • An OpenAI win on the merits, which would push publishers toward technical blocking, paywalled archives and legislative lobbying.
  • A muddled outcome, which would prolong the current era of case-by-case deals struck in the shadow of uncertainty.

In every scenario, verified provenance of training data becomes a commercial asset. The sanctions motion, whatever its result, has made 'what exactly did you train on, and can you prove it' the central question of the AI-media economy.

What should happen

The court should resolve the sanctions question quickly and transparently, because ambiguity here poisons an entire industry's negotiations. If OpenAI misled the court, the penalty should be proportionate and public. If it did not, the publishers should absorb that finding and litigate the merits without theatre.

Beyond this case, the lesson is legislative. A functioning market between AI firms and news producers needs disclosure standards that do not depend on subpoenas: auditable training-data records, standardised opt-outs, and privacy-preserving mechanisms for verifying model outputs. Courts are writing AI policy one discovery dispute at a time. That is the worst possible way to design the information economy, except for the current alternative, which is not designing it at all. Journalism can survive AI. It cannot survive a world where nobody can check what the machines were fed.

Sources

  • Washington Post - News outlets urge a judge to sanction OpenAI in a high-stakes AI copyright fight (9 July 2026)
  • TechCrunch - New York Times says OpenAI hid evidence in ChatGPT copyright trial (9 July 2026)
  • Bloomberg Law - New York Times seeks sanctions against OpenAI in copyright case (July 2026)
  • US News / Reuters - New York Times-led group asks court to sanction OpenAI in US copyright dispute (9 July 2026)
Share

You may also like to read