Microsoft says Copilot rarely reproduces publishers’ work [1]
Microsoft says fewer than 1 percent of 8.2 million selected Copilot conversations contained at least 16 words matching news content used to ground the AI model.[1] It presented the analysis while seeking summary judgment in consolidated copyright litigation brought by publishers and authors, but Th…
Microsoft says fewer than 1 percent of 8.2 million selected Copilot conversations contained at least 16 words matching news content used to ground the AI model.[1] It presented the analysis while seeking summary judgment in consolidated copyright litigation brought by publishers and authors, but The New York Times disputes Microsoft’s conclusions and alleges that Microsoft and OpenAI built competing commercial products using its journalism.[1]
Why it matters: Microsoft argues that the low rate of matching output supports treating LLM training as transformative fair use, while the plaintiffs contend that unauthorized use during training harmed their businesses regardless of how often Copilot reproduces text.[1]
Key insights: The analyzed logs were selected because they contained keywords associated with the news plaintiffs’ websites, making them comparatively likely to include the plaintiffs’ work.[1] | Microsoft says 59,545 conversations contained at least 16 words matching grounded news content, while a Center for Investigative Reporting expert identified 51 instances of substantial overlap with CIR work.[1] | In the authors’ case, Microsoft says only 24 responses contained at least 30 matching words and only 10 of 212 evaluated books produced any matches.[1] | The Trump administration filed a statement of interest supporting OpenAI in The New York Times case.[1]
Cheatsheet facts: What changed: Microsoft submitted chat-log evidence arguing that substantive reproduction by Copilot is rare.[1] | Why now: Microsoft is asking the judge for summary judgment that would end the case at an early stage.[1] | Watch next: Watch whether the judge grants summary judgment or allows the publishers’ and authors’ claims to continue in court.[1]
![Visual Cheatsheet Version A for Microsoft says Copilot rarely reproduces publishers’ work [1]. Full text follows for assistive technology.](https://keldura.ai/daily/ai-technology/2026-09-05/stories/1/cheatsheet.png?v=d935c04ff914f5daacc9d8adf0210b650b882dedd6e3d518195e2bf246a60c18)
Microsoft says fewer than 1 percent of 8.2 million selected Copilot conversations contained at least 16 words matching news content used to ground the AI model.[1] It presented the analysis while seeking summary judgment in consolidated copyright litigation brought by publishers and authors, but The New York Times disputes Microsoft’s conclusions and alleges that Microsoft and OpenAI built competing commercial products using its journalism.[1]
Why it matters: Microsoft argues that the low rate of matching output supports treating LLM training as transformative fair use, while the plaintiffs contend that unauthorized use during training harmed their businesses regardless of how often Copilot reproduces text.[1]
Key insights: The analyzed logs were selected because they contained keywords associated with the news plaintiffs’ websites, making them comparatively likely to include the plaintiffs’ work.[1] | Microsoft says 59,545 conversations contained at least 16 words matching grounded news content, while a Center for Investigative Reporting expert identified 51 instances of substantial overlap with CIR work.[1] | In the authors’ case, Microsoft says only 24 responses contained at least 30 matching words and only 10 of 212 evaluated books produced any matches.[1] | The Trump administration filed a statement of interest supporting OpenAI in The New York Times case.[1]
Cheatsheet facts: What changed: Microsoft submitted chat-log evidence arguing that substantive reproduction by Copilot is rare.[1] | Why now: Microsoft is asking the judge for summary judgment that would end the case at an early stage.[1] | Watch next: Watch whether the judge grants summary judgment or allows the publishers’ and authors’ claims to continue in court.[1]