The New York Times Case Against OpenAI Forced Disclosure of 20 Million ChatGPT Conversations
The New York Times sued OpenAI and Microsoft in December 2023, alleging that millions of its articles were used without permission to train generative models. The case survived motions to dismiss and moved into discovery, and in January a judge affirmed an order compelling OpenAI to produce a sample of twenty million ChatGPT user conversations.
The copyright question is what everyone is watching. The discovery order is what will matter longest.
What twenty million conversations means
The Times needs to demonstrate that OpenAI’s product reproduces its work and substitutes for it. You can’t prove that from the outside, because the outputs are generated per user, per session, and nobody keeps a public record. So the plaintiff asks for the logs.
That’s ordinary litigation logic and an extraordinary outcome. Twenty million conversations is a corpus of private exchanges between people and a system they treated as a place to think out loud. Medical worries, legal problems, work they didn’t want colleagues to see, drafts of things never sent.
The users aren’t parties. They can’t object, negotiate scope or know whether their session is in the sample. Protective orders and anonymisation apply, and both are meaningful and both have limits, because conversational text is self-identifying in a way structured data isn’t. People name their employer in the second sentence.
The precedent nobody argued for
Once a court orders one AI company to produce a conversation sample, the mechanism exists. Every subsequent plaintiff in every subsequent case now has a template, and there are more than forty active cases between AI companies and rights holders.
That creates an odd alignment. AI companies, who have every commercial reason to retain logs, suddenly acquire a legal reason to keep less. Litigation exposure is the one argument that has ever reliably shortened a retention policy in a technology company, and it works when privacy advocacy doesn’t, because it shows up on a risk register with a number attached.
Expect retention windows to get shorter, expect more aggressive de-identification, and expect it to be announced as a privacy improvement.
The case itself
The Times is reportedly seeking damages in the billions, and the case is widely expected to become the test precedent for how fair use applies to model training. That expectation may be optimistic. Cases this size settle, and a settlement produces a commercial agreement instead of a doctrine.
Look at what the Times has done alongside the litigation. It licensed content to Amazon in a deal reported at twenty to twenty-five million dollars a year, covering the paper, its cooking product and The Athletic for use in Amazon’s assistants and foundation models. That deal notably does not place Times content inside the major chat assistants or Google’s answer summaries.
So the Times is licensing where the terms suit it and suing where they don’t, which is a rational strategy and tells you the company doesn’t view AI access as a matter of principle. It views it as a market where the price hasn’t been set.
What a settlement would leave behind
If it settles, the industry gets a number instead of a rule. A number that a very large publisher with a very large legal budget managed to extract, which smaller publishers can cite and cannot obtain.
The doctrine stays unresolved, the next case starts over, and everyone below the top tier keeps negotiating in the dark.
That’s the likely outcome, and it’s why the discovery order may outlive the verdict.