The most closely watched IP dispute of the generative-AI era will decide whether unlicensed model training is fair use. Inside the consolidated In re: OpenAI proceeding and its landmark discovery fights.
In re: OpenAI, Inc. Copyright Infringement Litigation — U.S. District Court for the Southern District of New York (Hon. Sidney H. Stein; Mag. J. Ona T. Wang)
The most closely watched intellectual-property dispute of the generative-AI era remains unresolved, and its trajectory will likely determine the economics of model training for a decade. The Times filed suit in December 2023, alleging that OpenAI and its investor Microsoft ingested millions of its articles to train large language models without license or compensation, and that the resulting systems can regurgitate protected expression verbatim. The case has since absorbed related actions by other publishers and authors into a consolidated proceeding before Judge Stein, making it the central vehicle through which the federal courts will confront whether unlicensed training constitutes fair use under 17 U.S.C. § 107.
Procedurally, the litigation cleared its first major hurdle in the spring of 2025, when Judge Stein declined to dismiss the core infringement and contributory-infringement claims while paring back several ancillary theories. That ruling preserved the plaintiffs' central contention—that copying at the scale of model training is not excused by the transformative-use doctrine—and set the parties on a discovery path of extraordinary breadth. No trial date has been set, and the fair-use question that animates the case has not yet been decided on the merits.
Discovery itself has produced some of the most significant rulings to date. The plaintiffs sought production of a large sample of ChatGPT conversation logs to test how often the model reproduces copyrighted text, and Magistrate Judge Wang ordered production of a twenty-million-log sample over OpenAI's objections. In January 2026, Judge Stein affirmed that order, rejecting the argument that the magistrate's decision was clearly erroneous and finding that user-privacy interests were adequately protected by the discovery safeguards in place. The dispute over user logs previews the evidentiary fights to come, in which the frequency and fidelity of memorized output will bear directly on both infringement and the fourth fair-use factor.
The stakes are difficult to overstate. Statutory damages for willful infringement run to $150,000 per work, and the sheer number of asserted registrations means that an adverse verdict could translate into liability of a magnitude no defendant has faced in a copyright case. More important than any single award is the precedential question: whether the transformative-use logic that protected search indexing and snippet display extends to systems that learn from, and can reproduce, the expression they were trained on.
For practitioners, the case is a live laboratory for questions that will recur across every AI docket—how to define the relevant "use," how to measure market harm to a licensing market that is itself nascent, and how discovery obligations attach to training corpora and model weights. However the fair-use question is ultimately resolved, the record being built in the Southern District of New York will frame the analysis for every generative-AI dispute that follows.
Track this case with PacerPlus. In re: OpenAI is producing consequential discovery and fair-use rulings with no trial date yet in sight. Use pacerplus.com to monitor the docket in real time, get plain-English summaries of every new order and opinion, and ask questions about the record and the issues — then turn on PACEAlert to be notified the moment a ruling, filing, or scheduling change lands.