Policy

Microsoft says Copilot rarely reproduces news articles

Microsoft has filed legal documents arguing that its Copilot chatbot almost never regurgitates copyrighted text, a pivotal claim in its ongoing fair-use defense against major publishers.

The Verge AI3 days agoPolicy
Image: The Verge AI

In its latest push for a summary judgment, Microsoft submitted analysis of 8.2 million Copilot chat logs to demonstrate that the AI assistant rarely duplicates copyrighted material. The dataset, selected specifically for keywords related to the plaintiffs' publications, revealed that only 59,545 logs—fewer than 1 percent—contained at least 16 consecutive words in common with the news articles used to ground the model. Furthermore, an expert analyzing the data for the Center for Investigative Reporting identified just 51 instances of substantial overlap with their work.

The filings also addressed claims from book authors. An expert representing the authors found that the 8.2 million conversations yielded only 24 responses containing 30 or more matching words. Out of 212 books evaluated in the litigation, only 10 showed any matches at all. Microsoft argues these figures prove that Copilot does not serve as a substitute for the original works, reinforcing its stance that training large language models on copyrighted data constitutes transformative fair use.

Publishers strongly dispute this interpretation. Ian Crosby, lead counsel for The New York Times, stated that the discovery documents prove Microsoft and OpenAI "stole from The New York Times" to build competing commercial products. While the Times and other creators push for accountability, the legal battle has drawn federal attention, with the Trump administration recently filing a statement of interest supporting OpenAI in the New York Times case.

For AI practitioners and developers, the outcome of this litigation will define the boundaries of training data acquisition. If the court accepts Microsoft's argument that rare, incidental regurgitation does not undermine the transformative nature of LLMs, it will secure a major legal shield for commercial AI deployment. Conversely, a ruling against Microsoft could force developers to implement much stricter filtering systems or pay hefty licensing fees to avoid copyright liability.

This is our own summary of reporting by The Verge AI

More in Policy