Skip to content
News

Microsoft: Copilot Rarely Reproduced NYT Articles

Microsoft: Copilot Rarely Reproduced NYT Articles - Microsoft Copilot copyright
Microsoft says its Copilot chatbot rarely reproduced New York Times articles or books, in new filings defending its AI training against copyright claims.

Microsoft has told a US court that its Copilot chatbot rarely reproduces even full sentences from news articles and books, let alone substantial passages that could serve as a substitute for the original. The claim forms part of new legal filings submitted as the company defends itself against copyright claims brought by publishers, including The New York Times, and by book authors.

The dispute centres on allegations that Microsoft and OpenAI built commercial products on copyrighted works and now compete directly with the publishers and writers whose material was used. The company is seeking a summary judgement that would end the case at an early stage rather than proceeding to a full trial.

What the chat logs revealed

As part of the lawsuit’s discovery process, Microsoft provided 8.2 million Copilot chat logs to an expert hired by news publishers. According to the company, these logs were specifically selected because they hit on keywords implicating use of the publishers’ websites, making them the most likely to contain the plaintiffs’ works.

The resulting analysis, Microsoft says, showed that 59,545 of the conversations contained at least 16 words in common with news content used to ground the AI model. An expert for the Center for Investigative Reporting identified 51 instances of substantial overlap with the organisation’s work within the dataset.

In the parallel case brought by authors, an expert found that the 8.2 million conversations produced only 24 responses containing at least 30 matching words. Microsoft claims that just 10 of the 212 books evaluated returned any matches at all.

The fair use argument

Microsoft argues that the figures reinforce its position that using copyrighted content for AI training datasets should be regarded as fair use. While systems such as Copilot rely on copyrighted material, the company says the resulting products are used for purposes significantly different from the originals. The fact that they occasionally reproduce sections of text, it concludes, hardly undermines the transformative purpose of large language model training.

The New York Times rejected Microsoft’s conclusions. Lead counsel Ian Crosby said the documents and testimony uncovered during discovery led to only one conclusion, arguing that Microsoft and OpenAI had taken from the newspaper to make commercial products that substitute for its journalism, threaten its business and undermine its industry. He added that the publisher looked forward to the companies being held accountable.

The news publishers’ and book authors’ claims were consolidated under a single judge to streamline proceedings, despite objections from the plaintiffs. Microsoft submitted its filing on Friday. The Trump administration also filed a statement of interest in the New York Times case this week, supporting OpenAI. Should the judge side with the publishers and authors, the legal battle will continue in court.

Source
Image: theverge.com

The UK tech briefing

Smartphones, AI, computing and deals — the essential stories without the noise.

Mailing provider can be connected when your UK list is ready.

Shop on Amazon UK — Discover deals Shop on Amazon UK — Discover deals