US DOJ Backs OpenAI Against New York Times Lawsuit
The US Department of Justice has backed OpenAI and Microsoft in their copyright battle, arguing that training AI models on copyrighted data constitutes legally protected fair use.

The US Department of Justice (DOJ) has filed a brief siding with AI developers in the consolidated lawsuit brought by The New York Times against OpenAI and Microsoft. The Times sued the tech companies in late 2023, alleging that millions of its articles were used without authorization to train large language models like GPT-4. The newspaper is seeking billions of dollars in damages and the destruction of any models trained on its content. However, the DOJ argues that copying copyrighted text for the purpose of training AI does not constitute copyright infringement.
According to the DOJ, there is a clear legal distinction between copying data during the training phase and the final output generated by the model. The department notes that while entire works are copied during training, they are not made public, and the resulting outputs rarely show substantial similarity to the source texts. To illustrate this, the filing compares AI training to the author Joan Didion copying Ernest Hemingway's stories as a teenager to learn how to write. The DOJ warns that imposing broad liability for training would stifle the very creativity that copyright laws are designed to protect.
This stance directly challenges a previous report by the US Copyright Office, which rejected blanket fair use for AI training due to the unprecedented speed and scale of machine generation. The DOJ dismissed the assessment of former Copyright Register Shira Perlmutter, who was recently fired by the Trump administration, stating her report lacks binding legal authority and ignores established case law. The DOJ emphasized that fair-use determinations must be made on a case-by-case basis rather than through sweeping prohibitions.
For AI practitioners and developers, this intervention provides a significant boost. If the courts adopt the DOJ's reasoning, it will dramatically lower the legal risks of training frontier models on public internet data. Developers would not need to secure costly licenses for every piece of text ingested during training, provided the final outputs remain distinct from the training data. This could preserve the current open-web training pipeline and prevent a massive financial barrier to entry for smaller AI startups.
This is our own summary of reporting by The Decoder

