Microsoft says under 1% of Copilot logs echoed news content in NYT suit

Microsoft has filed a legal brief defending Copilot against copyright claims brought by The New York Times, other news publishers, and book authors, arguing that its chatbot rarely reproduces the copyrighted material it was trained on.
As part of discovery in the consolidated lawsuit, Microsoft handed over 8.2 million Copilot chat logs to an expert hired by the news publishers. Microsoft says these logs were deliberately selected because they hit on keywords tied to the plaintiffs' websites, making them the ones most likely to contain matching content. Even so, the resulting analysis found that only 59,545 of the 8.2 million logs, fewer than 1 percent, contained at least 16 words in common with the news content used to ground the underlying AI model.
Separate experts reached similarly small numbers. An expert for the Center for Investigative Reporting found just 51 instances of "substantial overlap" with CIR's work in the dataset. In the authors' lawsuit, an expert found only 24 of the 8.2 million Copilot conversations contained at least 30 matching words, and only 10 of 212 books evaluated had any matches at all.
The Times rejected Microsoft's framing. Its lead counsel, Ian Crosby, said in a statement that the discovery record leads to "only one conclusion: Microsoft and OpenAI stole from The New York Times to make commercial products that substitute for its journalism, threaten its business, and undermine its industry," adding that the Times looks forward to Microsoft and OpenAI "being held accountable for their theft." The Center for Investigative Reporting and the Authors Guild did not immediately respond to requests for comment.
Microsoft argues the low reproduction rate supports treating the use of copyrighted material for AI training as fair use. It says that while systems like Copilot rely on copyrighted content, they are used for purposes significantly different from the original works, and that occasional reproduction of text "hardly undermines the transformative purpose of LLM training."
The filing, submitted on Friday, asks the judge overseeing the consolidated case to issue a summary judgment that would end the lawsuit at an early stage. The publishers' and authors' claims were merged under one judge to streamline the litigation, over the plaintiffs' objections. The Trump administration separately filed a statement of interest in the New York Times case this week, in support of OpenAI. If the judge instead sides with the publishers and authors, the case will continue in court.
Key facts
- Microsoft's court filing says fewer than 1% of 8.2 million Copilot chat logs contained at least 16 words matching news content used to ground the AI model.
- Those 8.2 million logs were themselves pre-selected by keywords tied to the plaintiffs' websites, meaning they were the ones most likely to match, yet only 59,545 qualified.
- An expert for the Center for Investigative Reporting found 51 instances of substantial overlap; an authors'-suit expert found only 24 of 8.2 million conversations had 30+ matching words, and just 10 of 212 books evaluated had any matches.
- The New York Times' lead counsel Ian Crosby called the evidence proof of theft, while Microsoft says the numbers support a fair use defense for AI training.
- Microsoft filed on Friday for summary judgment to end the case early; the Trump administration separately filed a statement of interest supporting OpenAI.
Why it matters
The case tests whether training AI models on copyrighted news articles and books, and occasionally reproducing fragments of that text in chatbot answers, counts as fair use under copyright law. Microsoft is trying to turn a data point, how rarely Copilot's output actually overlaps with the source material, into the legal argument that decides the question. A summary judgment win for Microsoft would end this lawsuit before a jury ever weighs the fair use question directly; a loss would send the dispute over Copilot's use of news and book content to trial.
Who it affects
The direct parties are Microsoft and OpenAI as defendants, and The New York Times, other unnamed news publishers, the Center for Investigative Reporting, the Authors Guild and other book authors as plaintiffs. The judge overseeing the consolidated case will decide whether the lawsuit proceeds. The outcome also matters to any publisher or author considering a similar claim against an AI company, since it will show how courts weigh output-level reproduction rates as evidence in a fair use defense.
How to use it
Microsoft's filing offers a template other AI companies facing similar claims could point to: measure how often a model's actual outputs reproduce lengthy verbatim strings from copyrighted works, using the plaintiffs' own keyword criteria to select the sample, rather than arguing only about what went into training. Microsoft used that data to seek summary judgment, asking the judge to decide the fair use question as a matter of law without a trial. If granted, the case ends at this stage; if denied, the dispute over Copilot's use of news and book content continues in court.
How solid is it
The reported figures come from Microsoft's own legal filing and from experts hired by the opposing parties, not from an independent audit described in the article. The 8.2 million logs were themselves already selected using keywords tied to the plaintiffs' websites, specifically to surface the logs most likely to contain matching content, so the sub-1% match rate is measured against a sample stacked toward positive results, not a random slice of all Copilot traffic. The New York Times' lead counsel, Ian Crosby, explicitly rejects Microsoft's reading of the same evidence, calling it proof of theft rather than fair use. The Center for Investigative Reporting and the Authors Guild had not commented by publication, and no court has yet ruled on the summary judgment motion.
Risks and caveats
The article gives no calendar date for the filing, only that it was submitted on a Friday, and does not identify which New York Times articles or which of the 212 evaluated books were among the matched content. There is no timeline for when the judge will rule on summary judgment, and the article does not describe what legal weight the Trump administration's statement of interest carries beyond noting that it supports OpenAI. Most importantly, nothing here is decided: the case is still pending, and Microsoft's fair use argument has been neither accepted nor rejected by any court.
“The documents and testimony uncovered during discovery lead to only one conclusion: Microsoft and OpenAI stole from The New York Times to make commercial products that substitute for its journalism, threaten its business, and undermine its industry”
— Ian Crosby, The New York Times' lead counsel