Microsoft wants a federal court to know that its AI chatbot is not a plagiarism machine. In new legal filings responding to copyright claims from The New York Times and other publishers, the company argues that Copilot rarely reproduces even full sentences from news articles, let alone chunks substantial enough to substitute for the original work.
The argument rests on a massive trove of data: roughly 8.2 million Copilot chat logs that Microsoft handed over to an expert hired by the news publishers during discovery. And the company says these logs were deliberately selected to be the most damning ones possible.
The numbers behind the defense
Microsoft’s legal strategy here is essentially statistical. The company claims the 8.2 million logs were filtered specifically because they contained keywords implicating use of the news plaintiffs’ websites. In other words, these weren’t randomly pulled conversations about weekend dinner recipes. They were the outputs most likely to contain reproduced content from the publishers suing Microsoft.
Out of that filtered pool, Microsoft says the analysis identified 59,545 instances where Copilot’s output overlapped with plaintiffs’ content. That’s roughly 0.7% of the total sample, a number Microsoft clearly considers exculpatory.









