Wire flash
TechUS DOJ: AI training on copyrighted works generally qualifies as fair use
Editorial responsibility
- No named human review is recorded for this page.
- Source reporting is collected, normalized, translated or condensed automatically when needed.
- Automatically published source-backed update
The U.S. Department of Justice has formally intervened in the copyright lawsuit between OpenAI and The New York Times, filing a statement of interest that argues training large language models on copyrighted texts should generally qualify as fair use. This marks Washington's first formal intervention in the wave of AI training copyright lawsuits, though the filing is advisory and not binding on the court. The DOJ's argument distinguishes between acquiring material, training on it, and generating outputs, focusing specifically on the training stage. It contends that training serves a different purpose from publishing an article, as an LLM uses text to learn statistical relationships and generate new responses. Regarding market harm, the DOJ states that training itself does not substitute for the original work, and later AI-generated competition should not retroactively make the training unlawful. The administration also warned that blanket licensing requirements could disadvantage smaller AI companies and U.S. developers against foreign competitors. The court must still decide fair use on a case-by-case basis.
Source report
The U.S. government has formally endorsed OpenAI's position that training artificial intelligence on copyrighted text constitutes fair use, marking a significant legal win for the company and the broader AI industry.
Government's First Formal Intervention
In its first official intervention in the wave of copyright lawsuits surrounding AI training, the U.S. Department of Justice filed a statement of interest arguing that training large language models (LLMs) on copyrighted materials should generally qualify as fair use. While the filing is advisory and not binding on the court, it represents Washington's clearest stance yet on the issue.
Key Arguments from the Justice Department
The DOJ's filing distinguishes between three separate stages of AI development:
- Acquiring training material
- Training the model on that material
- Generating outputs
The department focuses its argument specifically on the copying that occurs during the training stage.
Purpose and Market Harm
The Justice Department argues that training serves a fundamentally different purpose from publishing an article. An LLM uses text to learn statistical relationships and generate new responses, rather than reproducing the original work.
On the question of market harm, the DOJ states that training itself does not substitute for the original work. Consequently, later AI-generated competition should not automatically render the earlier training unlawful.
National Security Considerations
The administration's backing of OpenAI's position is partly grounded in national security concerns. The filing warns that blanket licensing requirements could:
- Raise barriers for smaller AI companies
- Put U.S. developers at a disadvantage against foreign competitors
Implications for Ongoing Litigation
While courts must still decide fair use on a case-by-case basis, adopting the DOJ's framework would shift much of the legal pressure from model training toward data acquisition and specific outputs. The distinction leaves separate copyright questions unresolved, including:
- How training data was acquired
- Whether particular outputs reproduce protected passages
Source
rohanpaul_aiNeutral / independent
Part of this Story
U.S. DOJ backs OpenAI, argues AI training on copyrighted text is fair use in New York Times lawsuit