An eagle-engraved brass seal poised above two documents on a mahogany desk, the lower document crumpled, lit by a brass banker's lamp at left

The DOJ Said Training Doesn’t Copy. OpenAI Knew It Does.

/ Maxim Starkweather / 7 min read

On September 2, the Justice Department filed a statement of interest before Judge Sidney Stein in the Southern District of New York, formally backing OpenAI in its copyright fight with The New York Times. The government called AI training “extraordinarily transformative” and argued, specifically, that large language models “do not reproduce the text they train on.” This is the first time the federal government has formally stated its position on whether training a machine learning model on copyrighted content is legal. The conclusion it reached is probably right. The factual claim it rested that conclusion on is the one OpenAI cannot defend with a straight face.

What Washington Actually Filed

A statement of interest is not a party brief. The DOJ doesn’t join the case, doesn’t argue before the judge, and doesn’t control the outcome — Judge Stein decides independently. What a statement of interest does is signal where the executive branch stands, which matters in two ways: it tells the court this is a national priority, and it tells every other court, every other case, and every other lab what the government’s position is. This is not a narrow intervention in a narrow dispute between a newspaper and a software company. It’s the administration announcing what it thinks the law says about the entire category of machine learning.

The core argument is a fair-use claim under 17 U.S.C. § 107, the four-factor test. Purpose and character of use: training is transformative because LLMs build “generalized reasoning and language skills” rather than reproducing source material. The market effect: if you’re not reproducing the work, you’re not substituting for it. The government is not carefully arguing all four factors — it’s arguing the first one hard enough that the others don’t need to be answered. The filing calls training “extraordinarily transformative,” and that modifier is doing a lot of work.

The national security frame is explicit. The brief states: “The United States has a strong interest in continuing to develop a robust and competitive artificial intelligence industry that sets the standard for the practice and procedure of AI use globally.” It warns that restricting AI training under fair-use doctrine would “severely hamper ‘the Progress of Science and useful Arts.’” This is not new rhetoric — it’s the same frame the administration has used against export controls, against algorithmic disclosure requirements, against the EU’s AI Act. What’s new is the context: the DOJ is now attaching it to a specific legal theory, in a specific federal court, on the record.

The brief is broad in scope. It’s written to benefit not just OpenAI but any lab training on copyrighted text. Music labels, book publishers, news organizations — the government’s position now covers all of it. The filing’s outcome, if it persuades Stein, sets a template that extends well beyond this case.

The Hole in the Argument

The government’s case rests on one specific factual claim: large language models do not reproduce the text they train on. That’s the mechanism the brief uses to satisfy the market-substitution factor — if you’re not reproducing, there’s no market displacement. It’s also the one factual claim in the entire filing that the NYT’s lawyers have evidence against.

A mechanical counter ticking up inside a darkened server aisle, amber warmth against cold blue shadows

In July 2026, the Times alleged in court that OpenAI had concealed several things from discovery. Among them: a database of approximately 78 million de-identified ChatGPT conversations that OpenAI assembled internally to assess how much it was infringing on others’ work. And a set of internal tools called Project Giraffe that included a Bloom filter — a probabilistic data structure designed to detect and log instances where ChatGPT’s outputs regurgitated training data verbatim. Project Giraffe was built shortly after the lawsuit was filed. OpenAI subsequently deleted billions of ChatGPT outputs after the lawsuit commenced, which the court noted as a violation of its preservation orders.

OpenAI’s lawyers have argued that reproducing articles verbatim is “not what it is designed to do and not what it does.” That defense is harder to sustain when you’ve simultaneously built and concealed a system designed specifically to monitor when it does. The Times’ lead counsel put it directly: “If OpenAI genuinely believed that copying our clients’ journalism was fair and legal, it wouldn’t have hid the truth about having done it.”

To be precise about what Project Giraffe proves: it proves OpenAI thought reproduction happened often enough to be worth monitoring at scale. It doesn’t prove any specific output reproduced any specific article. OpenAI could argue that regurgitation detection was a defensive precaution, that engineers built Giraffe to find edge cases, that the Bloom filter logged very few hits. What it cannot argue cleanly is that the product behaves as if reproduction doesn’t happen — because that’s the behavior that makes a regurgitation detector unnecessary. You don’t instrument for a phenomenon you’re certain doesn’t exist.

The DOJ just filed a brief in the same case saying LLMs don’t reproduce text. The NYT is about to ask what Project Giraffe was for.

The Argument the Government Could Have Made

Fair use under § 107 is not a zero-reproduction test. Factor one — purpose and character of use — asks whether the use is transformative. The Supreme Court established in Campbell v. Acuff-Rose (1994) that transformativeness is the key consideration, and that transformation doesn’t require zero reproduction of source material. You can take substantial portions of a work and still prevail on fair use if the purpose is genuinely transformative and the market effect is distinct.

Shipping boxes crossing from a tungsten-lit loading dock into a blue-glowing server room beyond

The government could have argued: training is transformative even when it occasionally results in memorization, because the purpose of training is to build a generalized model, not to reproduce any specific work; the market for LLM outputs is different from the market for newspaper subscriptions; and verbatim regurgitation in outputs is a defect, not the purpose of the system. That argument doesn’t require the factually contested premise that LLMs “do not reproduce the text.” It would survive Project Giraffe on the stand. It’s also the stronger argument on the merits — transformation is about intent and function, not about whether any reproduction ever occurs.

Instead, the DOJ chose the categorical denial: doesn’t reproduce. That’s a cleaner sentence to put in a brief. It’s also a hostage to fortune. If Judge Stein asks about Project Giraffe — and the NYT’s lawyers have given him every reason to — the government’s brief has just helped frame the question. You said it doesn’t reproduce. Here’s the detector they built for when it does.

There’s a version of this outcome where the brief still achieves its goal: Judge Stein could find for OpenAI on transformativeness regardless of the reproduction question, treating the DOJ’s factual claim as advocacy rather than evidence. Courts don’t have to accept the government’s factual framings when ruling on fair use — the legal argument can survive even if the supporting factual claim is contested. But building a legal doctrine on a factual premise that the same case’s evidence challenges is how doctrine gets built badly, and the template this case sets will govern AI copyright disputes for a decade.

What the Brief Leaves Unaddressed

The DOJ’s filing addresses one question: whether the training itself is fair use. It says nothing about how labs obtained their training data. That’s a different question, and it’s the one where liability has already materialized.

Anthropic’s settlement with book publishers earlier this year — $1.5 billion over the Project Panama claims — was about acquisition, not training. Anthropic was alleged to have obtained unauthorized copies of books for use in training data. The issue wasn’t whether training on those books was legal. It was whether receiving and storing unauthorized copies constituted infringement, regardless of what you did with them afterward. The DOJ brief would not have changed that outcome. Training being fair use doesn’t immunize a lab from liability that arose before training began.

The music cases now pending — Sony Music and Warner Chappell’s suits against Anthropic, similar actions against other labs — mix acquisition and training in ways courts haven’t fully sorted. The DOJ’s position is that at least one half of each case has a clean answer: the training phase is fair use. The acquisition question — how you got the recordings, what licenses existed, what you agreed to in the terms of whatever service you were scraping — is what remains.

The government made the right call on AI training. The argument it made for that call rests on the one piece of the case OpenAI’s own evidence undermines. Project Giraffe was a regurgitation detector. OpenAI built it, ran it, and allegedly hid it from discovery in a case where the government just argued regurgitation doesn’t happen. That discrepancy has to be resolved before Stein rules. Whoever resolves it — the court, a remedial finding, or a settlement — will be setting the factual floor under the largest copyright doctrine shift in a generation. The foundation deserves better than this.

An eagle-engraved brass seal poised above two documents on a mahogany desk, the lower document crumpled, lit by a brass banker's lamp at left

AI-generated editorial illustration · TemperatureZero · September 3, 2026

Keep reading the signal

Get the Daily Signal — a concise briefing on what actually matters in AI and the systems around it.

Subscribe Free

Continue the archive

Latest BriefingsArticlesAbout Temperature Zero