Open Editor
Privacy
8 min read

NYT vs. OpenAI: Both Sides Just Asked the Judge to End the Case — On Opposite Grounds

Three sets of lawyers filed for summary judgment on the same day, describing the same evidence in almost incompatible terms. Here's what each side's own numbers actually say.

FiledWhat Happened
Sept 3, 2026DOJ files statement supporting OpenAI's fair use defense
Sept 4, 2026OpenAI, Microsoft, and five publishers all move for summary judgment
Not yet setOral argument requested by both defendants; no ruling date

The New York Times sued OpenAI and Microsoft in December 2023. Nearly three years and one consolidation with four other publishers later, the case reached a real inflection point on September 4, 2026, when OpenAI, Microsoft, and the news plaintiffs each filed for summary judgment in the same Southern District of New York proceeding, asking Judge Sidney H. Stein to resolve the central legal questions before any jury gets involved. The filings, part of the consolidated multidistrict litigation, don't just disagree on the law — they describe the underlying facts in terms that barely overlap.

Who's Actually in the Case Now

What started as a Times-only suit is now five plaintiffs deep: The New York Times Company, a group of publishers including the New York Daily News, Chicago Tribune, and Denver Post, plus Ziff Davis (owner of CNET, PCMag, and Mashable), the Center for Investigative Reporting, and The Intercept. Together they assert more than 10.8 million works were used without authorization. OpenAI's brief counters that the plaintiffs' own experts identified only about 6.3 million of those as actually used in training, leaving roughly 4.5 million works OpenAI says have no evidence of copying behind them at all.

The Regurgitation Numbers: A Rare Point of Rough Agreement

The most concrete factual dispute involves how often ChatGPT reproduces publisher text verbatim — and here the two sides' own figures land closer together than the rest of the case would suggest. OpenAI's expert searched a sample of 20 million ChatGPT conversation logs, produced through an earlier discovery order, and found 24 instances of verbatim regurgitation. The publishers' own experts, examining the same category of material, arrived at similarly small rates across the different plaintiffs.

PlaintiffVerbatim Regurgitation Rate
OpenAI's overall sample0.00012% (24 of 20M logs)
NYT / Daily News Plaintiffs0.00011%
Ziff Davis0.00002%
Center for Investigative Reporting0.000008%
The Intercept0.000002%

Where the sides diverge sharply is on what those numbers mean. OpenAI frames them as proof its models don't function as substitutes for the original journalism. The publishers ran a separate, much larger set of tests — reportedly more than 250 million prompts using non-public model configurations — and found partial regurgitation in under 2.6% of targeted works, extracting an average of under 4.4% of each one's content. OpenAI's brief places that figure well below the 16% extraction rate that was found sufficient for fair use in the Google Books litigation.

Did Silence on robots.txt Count as Consent?

One of OpenAI's more pointed arguments concerns Browse, the web-browsing feature ChatGPT uses to answer questions about current events. OpenAI disclosed its browsing crawler and stated it would honor robots.txt exclusions in a March 2023 blog post. According to OpenAI's filing, Times employees discussed that announcement internally within nine days — but the Times reportedly didn't implement a hard block for the crawler until April 2024, more than a year later. OpenAI's position is that any browsing-related copies made in that window were impliedly licensed, since a two-line addition to a public text file would have refused the crawler outright. A footnote in the filing cites Reuters Institute research finding that by the end of 2023, roughly half of the most-visited news sites across ten countries had already blocked OpenAI's crawlers — meaning many publishers acted quickly while the Times, on this account, did not.

There's a real tension worth naming here: this same "publishing something publicly isn't the same as consenting to its reuse" argument is close to the opposite of what EU regulators concluded about AI training data, as covered in the earlier post on the EDPB's web scraping guidance. US litigation and EU regulatory guidance are, on this specific question, pulling in different directions.

The Traffic Argument: Revenue Is Up, But So Is the Crawl Ratio

Both sides lean on the fourth fair-use factor — market harm — and both have retained economists to make the case. OpenAI points to its own expert's finding of no measurable drop in visits to publisher sites tied to ChatGPT adoption, alongside the fact that Times digital subscriptions passed 12 million in 2025 and Times second-quarter digital advertising revenue for 2026 rose to roughly $114 million, up about 21% year over year. The publishers respond that fair use is an affirmative defense the defendants have to prove, not something the plaintiffs have to disprove, and point to a different metric entirely: crawl-to-referral ratios. Citing Cloudflare data, the publishers' brief states that by mid-2025 Google was crawling roughly 18 pages for every visitor it referred back to a site, while OpenAI's ratio stood at about 1,500 to 1 — and a Cloudflare dashboard opened in mid-2026 has since recorded ratios as high as nearly 50,000 to 1 for some sites.

The "Pink Slime" Argument

Separately from the core copyright claims, the publishers' brief devotes real attention to what it calls low-quality, AI-generated content published to resemble legitimate local journalism. Using OpenAI's own published API pricing, the publishers calculate that generating a million 500-word news-style articles would cost roughly $6,800 and require no reporter, editor, or fact-checker — and they cite an operation running roughly 200 AI-generated sites posing as local newsrooms as a real-world example of that math playing out. The legal theory attached to this argument, market dilution, comes from a federal judge's reasoning in a separate AI copyright case involving Meta, and holds that news markets may be especially vulnerable to being crowded out by cheaply generated competing content, distinct from direct copying.

Where This Sits Among Other AI Copyright Cases

The Times case isn't happening in isolation, and the outcomes elsewhere cut in both directions. Anthropic agreed to a $1.5 billion settlement in a separate authors' copyright case in September 2025 — the largest publicly reported figure in this category of litigation. Meta, by contrast, won summary judgment on fair use grounds in a different case in June 2025, though on a factual record the presiding judge described as thin. Those two outcomes give both sides in the Times case a precedent to point to, and neither one settles how Judge Stein is likely to rule here.

CaseOutcome
Anthropic (authors' case)$1.5B settlement, Sept 2025
Meta (separate case)Won summary judgment on fair use, June 2025
NYT v. OpenAI / MicrosoftCross-motions for summary judgment pending, filed Sept 2026

What Happens Next

Both OpenAI and Microsoft have requested oral argument, and none has been scheduled yet. No trial date exists. The range of outcomes is genuinely wide: the court could grant summary judgment to one side on liability or fair use and send only damages to a jury, could reject both sides' motions and send the whole case to trial, or — as has happened in comparable cases — the parties could settle before any of that plays out. The discovery fight over ChatGPT logs that produced the regurgitation-rate evidence, covered in more detail in the earlier post on how AI chatbot conversations are being used as courtroom evidence, is itself a reminder of how much of this case has turned on data nobody expected to become part of a public record.


The honest summary of where this stands: both sides filed the strongest version of their case on the same day, their own regurgitation-rate numbers are closer together than either side's framing suggests, and the rest of the dispute — traffic impact, implied licensing, market dilution — remains genuinely contested with real evidence on both sides. No ruling is imminent, and whichever way Judge Stein eventually decides, the outcome is expected to set a reference point for AI licensing negotiations well beyond this one case.

For questions or inquiries contact us at info@cleartexteditor.com