Microsoft says Copilot rarely reproduces publishers’ work [1]
Microsoft says fewer than 1 percent of 8.2 million selected Copilot conversations contained at least 16 words matching news content used to ground the AI model.[1] It presented the analysis while seeking summary judgment in consolidated copyright litigation brought by publishers and authors, but The New York Times disputes Microsoft’s conclusions and alleges that Microsoft and OpenAI built competing commercial products using its journalism.[1]
Microsoft argues that the low rate of matching output supports treating LLM training as transformative fair use, while the plaintiffs contend that unauthorized use during training harmed their businesses regardless of how often Copilot reproduces text.[1]
Key insights
- The analyzed logs were selected because they contained keywords associated with the news plaintiffs’ websites, making them comparatively likely to include the plaintiffs’ work.[1]
- Microsoft says 59,545 conversations contained at least 16 words matching grounded news content, while a Center for Investigative Reporting expert identified 51 instances of substantial overlap with CIR work.[1]
- In the authors’ case, Microsoft says only 24 responses contained at least 30 matching words and only 10 of 212 evaluated books produced any matches.[1]
- The Trump administration filed a statement of interest supporting OpenAI in The New York Times case.[1]