
Every large language model now in commercial use was built by ingesting an ocean of text its creators did not license — and the U.S. government has now told a federal court, in the plainest terms yet, that this is exactly how the law should work.
Key Points
- The Justice Department filed a formal Statement of Interest in the Southern District of New York backing OpenAI’s fair-use defense against The New York Times, marking the first time the federal government has weighed in on AI copyright litigation.
- The DOJ’s core legal move is separating “training” from “output” — arguing a model learning statistical patterns from text is a fundamentally different act, legally, than a chatbot reproducing that text.
- The Times’ case rests on the opposite premise: that the models can be prompted to return near-verbatim copies of its journalism and function as market substitutes for it.
- Courts have not yet settled the question uniformly; 2025-2026 rulings in related cases have split on how much weight to give market harm versus transformative purpose.
- The DOJ invokes national security and economic competitiveness as stakes, though the public record offers no disclosed empirical study quantifying those claims.
What the Justice Department Actually Argued
On September 1-2, 2026, the DOJ filed a Statement of Interest in In re OpenAI, Inc. Copyright Infringement Litigation, the consolidated Manhattan case brought by The New York Times against OpenAI and Microsoft. A Statement of Interest is a formal mechanism by which the executive branch tells a court it has a stake in how a private lawsuit is decided — not as a party, but as a voice claiming the outcome touches sovereign interests. The DOJ used it to say, in effect, that an adverse ruling against OpenAI would not just cost one company; it would unsettle the legal foundation the entire American AI industry has been built on. Reuters quoted the brief declaring “a strong interest in this court rejecting any argument that training LLMs on copyrighted texts violates copyright law,” citing scientific advancement and national security as the values at stake.
The brief’s central analytical move is a distinction that will likely outlive this particular case: training, it argues, “should be considered separately from a chatbot’s output”. Under this framing, when a model ingests a New York Times article during training, it is not copying the article’s expressive content — its prose style, its narrative choices, its specific sentences — but extracting statistical relationships among words, the way a linguist might tabulate sentence structures across a corpus without reproducing any single text. The DOJ describes this as using the work “not to duplicate the work’s expressive content, but as part of a process to learn and act on statistical patterns in written text”. That is a transformation-of-purpose argument, the doctrinal heart of fair use since the Supreme Court’s 1994 Campbell v. Acuff-Rose ruling, applied here to a technology Congress never contemplated.
Why the Times Disagrees, and Why Its Case Is Not Trivial
The Times’ lawsuit does not quarrel with the abstract proposition that transformative uses of copyrighted material can be lawful; it quarrels with whether that description fits what OpenAI’s models actually do. The paper’s complaint alleges the systems “can be coaxed or prompted to return verbatim or close to verbatim copies of copyrighted New York Times stories”, and its amended complaint sharpens this into a doctrinal counter-thesis: “because the outputs of Defendants’ GenAI models compete with and closely mimic the inputs used to train them, copying Times works for that purpose is not fair use”. This is not a vague grievance about technology outpacing law — it is a specific factual claim about output behavior that, if proven at scale, directly undercuts the DOJ’s training/output separation. If a model can regurgitate a Times article closely enough to substitute for reading the article itself, the transformative-purpose argument weakens considerably, because the fourth fair-use factor — effect on the market for the original — cuts hard against the defendant.
The Times has also been explicit that its objection is not to AI as such but to unlicensed commercial use of its journalism. Its general counsel said in 2023 that the paper recognized generative AI’s potential but believed “the use of our work to create GenAI tools must come with permission and an agreement that reflects the fair value of that work, as the law provides”. That framing — payment and permission, not prohibition — has proven durable; the Times later brought a similar theory against Perplexity, alleging the search company scraped its articles, videos, and podcasts to answer user queries without a license. The pattern across both suits is consistent: the Times is not arguing that machines shouldn’t read news, but that whoever profits from a machine trained on its work owes it something.
Where the Legal and Factual Record Actually Stands
It matters that the DOJ’s brief is a policy and doctrinal argument, not a forensic one. Nothing in the public record shows the department examined OpenAI’s actual training logs, memorization rates, or deduplication practices to confirm the “statistical patterns, not duplication” characterization holds for the specific datasets at issue in this case. The brief’s national-security and competitiveness claims rest on assertion rather than disclosed economic modeling. That is not unusual for a Statement of Interest — the DOJ is arguing law and policy, not serving as an expert witness — but it does mean the government’s intervention should be read as a doctrinal position with institutional weight, not as an independent fact-finding verdict on what OpenAI’s models actually do with a given Times article.
The broader judicial landscape offers a genuinely mixed signal, which is the honest state of the law right now. Federal courts in Bartz v. Anthropic and Kadrey v. Meta issued the first merits-stage rulings on AI training and fair use in 2025, and both leaned toward treating training as highly transformative while still scrutinizing the provenance of training copies — did the company license the text or pirate it — and market harm separately. The U.S. Copyright Office’s own 2025 report on generative AI training explicitly declined to issue a blanket rule, insisting the fair-use analysis is “context-specific” and will not yield one answer across all AI systems. That refusal to generalize is itself the most important fact in this entire dispute: neither Congress nor the courts have settled this, and every ruling to date has been fact-bound to its specific defendant, its specific dataset, and its specific evidence of output behavior.
The DOJ just backed OpenAI’s fair-use argument in its copyright battle with The New York Times.
The government argues AI training on copyrighted material can be transformative and that overly restrictive rules could hurt U.S. AI development.
But this is not a court ruling. pic.twitter.com/xl3EJ7DPlT
— Snaphomz_AI (@Snaphomz_AI) September 4, 2026
What This Means Going Forward
The DOJ’s intervention raises the stakes of the OpenAI-Times case well beyond the parties involved, because a ruling adopting the training/output distinction as broad federal policy would function as a template for the dozens of similar suits now working through courts against Anthropic, Meta, Perplexity, and Suno. Conversely, if the Times can put concrete, reproducible evidence of near-verbatim output in front of the judge — not hypothetical risk, but demonstrated substitution — the government’s clean doctrinal line gets harder to sustain, because copyright law has never protected a use simply because the intermediate step was clever; it protects a use because the end result doesn’t supplant the original’s market. The next meaningful developments will not come from further government statements but from the trial record itself: discovery into training data provenance, and any output audits that settle, empirically, whether these models are engines of learning or engines of reproduction dressed up as one.
Sources:
reason.com, nytimes.com, washingtonpost.com, insiderfinance.io, aichatdaily.com, wsj.com, cdn.arstechnica.net, x.com, lexology.com, arl.org, papers.ssrn.com
© fixthisnation.com 2026. All rights reserved.











