The US Department of Justice has sided with AI companies in the consolidated lawsuit involving the New York Times and other rights holders, arguing that training AI models on copyrighted material qualifies as fair use.
The New York Times sued OpenAI and Microsoft in late 2023 in a Manhattan federal court, alleging that millions of NYT articles were used without permission to train models like GPT-4 and build products that compete with the newspaper as an information source. The Times cited billions of dollars in damages and demanded the destruction of language models trained on its articles. The case has escalated considerably since then and is widely considered a bellwether for how courts will handle copyright and AI training.
The DOJ now says the copyrighted text used for LLM training in this case doesn't amount to copyright infringement. The distinction between training and output is what matters. During training, entire works are copied but never made publicly available, and the outputs "often if not always lack substantial similarity" to the originals. A blanket theory of market harm that conflates the two is legally wrong. Others disagree.
The filing invokes Joan Didion as an analogy. As a teenager, she copied Hemingway's stories to understand how his sentences worked. The DOJ argues that, under the logic of the Kadrey ruling, Didion could have faced liability whenever she published because her learning process and subsequent writing would have been treated as a single use. Citing an earlier ruling, the department contends that it would be unthinkable to require people to pay whenever they later draw on a book to write something new in a new way.










