blog

The United States told a federal judge that LLM training is fair use. Three of its nineteen numbered pages go after a ruling Meta won

On 1 September the United States filed twenty pages, nineteen of them numbered, in the consolidated OpenAI copyright cases arguing that copying a book or a news article in order to train a large language model is fair use. It states the position without hedging, twice. The heading over its central section reads "Training Of LLMs On Written Works Is Exceedingly Transformative," and the section says it again in plainer words: "In sum, the use of copies to train LLMs is extraordinarily transformative."

The docket is In re OpenAI, Inc. Copyright Infringement Litigation, 25-md-3143 (SHS) (OTW), in the Southern District of New York before Judge Sidney Stein. Associate Attorney General Stanley E. Woodward Jr. and Assistant Attorney General Brett Shumate head the signature block, and Senior Counsel Michael Weisbuch signed it. The caption page carries the MDL number. The ECF stamp on every page carries a member-case one, Case 1:25-cv-03483-SHS-OTW, Document 316, Filed 09/01/26.

The government is not a party to any of it. It appears under 28 U.S.C. § 517, which its own first footnote describes, quoting Gil v. Winn Dixie Stores, as a statute that "contains no time limitation and does not require the Court's leave." Read that the other way and you have the filing's actual weight. Nobody asked for it, nobody has to answer it, and Judge Stein may adopt every word or none. A statement of interest is a letter the court is free to read.

Its reach is another matter. The caption line says "This Document Relates To: All Matters," and footnote 11 makes the ambition explicit: the government's "legal arguments apply similarly to all parties in this litigation and the related cases, including book authors and publishers." It is written against the Times and aimed at everyone consolidated behind the Times.

It takes a position on all four factors, and two of them are in a footnote

IPWatchdog's account, which is representative, has the brief "centered on the first and fourth statutory fair-use factors." That is exactly where the two argument sections sit, Section C on purpose and character and Section D on market effect. It is not the whole brief. Footnote 16 handles factors two and three in a single paragraph, both in OpenAI's favour. The nature of the copyrighted work "likely supports a fair-use ruling" because the training use's "transformative purpose . . . inevitably involves the second factor as well." So does the amount used, because "[t]raining an LLM, in and of itself, does not make any copied works accessible to the public."

Four factors, then. Two of them disposed of in a footnote, which is its own kind of statement.

The sustained attack is on Kadrey, a case the AI company won

The brief's own pages 15 through 17 go after Kadrey v. Meta Platforms. Meta won that case. In June 2025 Judge Vince Chhabria granted Meta partial summary judgment on fair use for training, and then wrote at length about a theory the plaintiffs had barely argued: that LLM outputs could dilute the market for human writing so broadly that the fourth factor would sink fair use even where no single output resembles any original.

The brief's name for this is "[c]ontrary dicta in a Northern District of California decision." Its assessment is that the theory "misapplies copyright principles to LLM training" and is "deeply flawed," because the Kadrey court "improperly collapsed LLM training and LLM outputs into a single continuous use, then applied a capacious, genre-level understanding of 'substitution.'"

The illustration the government reaches for is Joan Didion, who as a teenager "would type out" Hemingway's stories "to learn how the sentences worked." By Kadrey's logic, the brief argues, "Didion should have incurred liability to Hemingway every time she published a piece, because the process by which she trained herself and the process by which she produced works was all one use."

Market dilution is the theory the plaintiffs would most like to rely on, precisely because it does not require proving that any particular output copies anything. Whoever files next in this MDL now has to answer three pages of federal government written specifically against it.

Footnote 17 goes past the case law. The Register of Copyrights' May 2025 report, Part 3 of Copyright and Artificial Intelligence, reached a conclusion close to Chhabria's. The brief's answer is that the Register "appeared to endorse a similar theory in a report," that "[h]er understanding does not warrant deference" under Loper Bright Enterprises v. Raimondo, and that her "threadbare reasoning ignored all the caselaw emphasizing the required use-by-use analysis."

The same footnote notes, in a subordinate clause, that the Register "is currently challenging her removal." Shira Perlmutter was removed on 10 May 2025 by Todd Blanche, then Deputy Attorney General, two days after Trump appointed him acting Librarian of Congress. She is still in office under a D.C. Circuit injunction the Supreme Court declined to disturb on 30 June 2026. So the department asking a court to disregard the Register's report is the department whose own deputy signed her removal, and the brief mentions the litigation only to note that it exists.

Training, not acquisition, and acquisition is where the money moved

The brief lays out three stages and then narrows to one. A developer first collects data at what it calls "the acquisition (or 'collection' or 'pre-training') stage," then trains on that data, then serves outputs. "The United States focuses on the question whether the use of copyrighted works at the training stage . . . constitutes fair use." Outputs get a paragraph conceding they raise separate questions, to be judged output by output. Acquisition gets that one appearance in the stage description and nothing afterwards.

That silence is the load-bearing part. Anthropic did not pay $1.5 billion because it trained Claude on books. It paid over how it obtained them. Judge Araceli Martínez-Olguín granted final approval to the Bartz v. Anthropic settlement on 20 July 2026, and the Authors Guild's account of the approval is specific about its shape. Class members "release only claims relating to Anthropic's past acquisition and copying of their works—the 'inputs' side—through August 25, 2025." And: "Claims based on AI outputs are not released, and neither are any claims of any kind about future conduct."

The brief cites Bartz four times and never once for the half that cost Anthropic the money. One of those citations is the line that the copying is "transformative — spectacularly so." Search the twenty pages for "settlement," for "LibGen," for "Library Genesis," and you get nothing. The one instance of the string "Pirate" is the name of a website the brief cites in a footnote about data-centre water use.

Two arguments that are policy, not doctrine

The first is national security. The brief quotes a 2022 GAO report warning that "[f]ailure to adopt and effectively integrate AI technology could hinder national security," and reaches its point in one sentence: "Rules of law that make it significantly more difficult to develop a robust AI industry in the United States therefore threaten national security and give a competitive advantage to foreign adversaries who are not so encumbered." No adversary is named anywhere in the document.

The second is competition, and it is pointed at the plaintiffs. If training requires licences, the brief argues, "only the largest technology companies might have the capital necessary to pay licensing fees," and those fees "would disproportionately benefit legacy media outlets due to the sheer volume of their written publications." The phrase it lands on is "large subsidies for old mainstream media companies." On that account a newspaper suing for a licensing regime is asking for an entry barrier only the biggest labs could clear.

The disclosure the filing does not make

On 2 July 2026 the Financial Times reported, and CNN, CNBC, Forbes and Reuters carried the same day, that OpenAI had discussed handing the federal government a roughly 5% equity stake, worth about $42.6 billion against the $852 billion valuation set in a March funding round. CNN calls the talks "early conversations" and says any deal "might require an act of Congress to implement." For the names, the Reuters wire: "Altman has discussed the stake sale with Trump, Commerce Secretary Howard Lutnick and Treasury Secretary Scott Bessent, the FT said." Nothing has been concluded. Above the Law raised it on 3 September as a conflict the filing does not acknowledge, and that framing is Above the Law's rather than a finding anyone has made on the record.

What is checkable is the document. It carries one disclaimer, footnote 2, and it is about something else entirely: that the government does not contend the conduct at issue was authorised by it or undertaken for its benefit under 28 U.S.C. § 1498. The words "equity" and "stake" do not appear in the twenty pages. Whether an unconsummated discussion of an ownership interest is the kind of thing a § 517 filing has to disclose is a question nobody has answered. It is not answered here.

Who this lands on

For the authors, publishers and papers consolidated in the MDL, nothing has been decided and one thing has changed: their strongest fourth-factor theory now has the United States written against it in a document their opponents will attach to everything, and their next brief has to deal with that before it deals with OpenAI.

For anyone building on scraped text, the split matters more than the headline does. A federal endorsement of training as fair use says nothing about where a corpus came from, and the $1.5 billion that changed hands in Bartz was about exactly that. Read the brief as an argument about one of three stages, because that is how it describes itself.

The Times answered on 2 September, saying the administration "is siding with a handful of trillion-dollar AI companies at the expense of the countless American creators whose work they stole." Graham James, a spokesperson for the paper, put the rest in an emailed statement: "Both AI and creators can thrive — AI companies simply need to pay fairly for the content that makes their products possible, as copyright law requires."

Sources

All retrieved on 5 September 2026. Quotations of the brief come from the PDF linked first.

  • Statement of Interest of the United States, In re OpenAI, Inc. Copyright Infringement Litigation, 25-md-3143 (SHS) (OTW) (S.D.N.Y. filed 1 September 2026), 20pp., ECF stamp Case 1:25-cv-03483-SHS-OTW Document 316 — storage.courtlistener.com
  • Authors Guild, "Court Grants Final Approval of $1.5 Billion Anthropic Copyright Settlement", on the 20 July 2026 approval in Bartz v. Anthropicauthorsguild.org
  • Associated Press via the Boston Globe, "Trump administration backs OpenAI in New York Times' copyright case over training of chatbots", 2 September 2026, carrying the Times statement — bostonglobe.com
  • Lowenstein Sandler, "U.S. Government Backs Fair Use for AI Training in OpenAI Copyright Litigation", 3 September 2026, describing the filing as "the federal government's first direct intervention" in the AI training cases and as "not binding on the court" — lowenstein.com
  • Above the Law, "DOJ Tells Court AI Training Is Fair Use, Forgets To Mention It's Negotiating A Stake In OpenAI", 3 September 2026 — abovethelaw.com
  • CNN Business, "OpenAI in talks to give Trump administration a 5% stake in the company, FT reports", 2 July 2026, source of the $42.6bn and $852bn figures, "early conversations" and the act-of-Congress line; it names neither secretary — cnn.com
  • Reuters wire, "OpenAI discussed giving 5% stake to Trump administration, media report says", 2 July 2026, read via The Globe and Mail, an account that names Lutnick and Bessent — theglobeandmail.com
  • Civil Rights Litigation Clearinghouse, Perlmutter v. Blanche, No. 1:25-cv-01659 (D.D.C.), on the 10 May 2025 removal and the 30 June 2026 Supreme Court denial — clearinghouse.net
  • IPWatchdog, Eileen McDermott, "DOJ Sides with OpenAI, Warns Obstacles to AI Development Threaten National Security", 3 September 2026 — ipwatchdog.com
  • Goodwin, "Northern District of California Judge Rules That Meta's Training of AI Models Is Fair Use", on the June 2025 Kadrey summary-judgment ruling and its market-dilution dicta — goodwinlaw.com

Primary evidence