Skip to content
News

AI Firms Knew Chatbots Threatened Journalism, Filing Shows

AI Firms Knew Chatbots Threatened Journalism, Filing Shows - AI chatbots journalism threat
AI firms knew chatbots posed an existential threat to journalism and used copyrighted text to train models, a newly unredacted court filing reveals.

AI companies used enormous quantities of published text to train their large language models, and executives at those firms were aware that this amounted to a theft of copyrighted material, according to comments cited by news publishers in a court filing unredacted this week.

The brief was initially lodged earlier this month in a proceeding that brings together several related copyright infringement cases against ChatGPT maker OpenAI and its partner Microsoft. Among the claimants are The New York Times and Ziff Davis, the owner of CNET. The full exhibits showing the context of the quotes remain under seal.

Microsoft’s director of applied science, Brent Hecht, described the practice as “an astonishing theft of unprecedented proportions” and possibly the “largest theft of labour in human history”, according to the publishers’ filing.

Fair Use Argument Under Scrutiny

The comments cited by the publishers appear to weaken the argument advanced by AI firms that the use of copyrighted material was legally protected as fair use. Under this doctrine, the use of copyrighted material may be protected depending on how it is used, the nature of the work, the amount used and the effect of that use on the market for the work.

Microsoft chief executive Satya Nadella said under oath that conversations with chatbots delivered information “right there on the website on the AI platform versus needing to go to the underlying source”, such as the website of the publisher that first reported it, according to the filing. An OpenAI executive similarly wrote that publishers faced an “existential threat” from products such as the company’s chatbot.

The brief also describes steps developers took to evade paywalls, including that of The New York Times. It recounts how an OpenAI employee told the company’s president, Greg Brockman, about “a hack to get around nytimes paywall”, to which Brockman replied, “ah nice.”

The filing states that Nadella testified that anything behind a paywall “should be licensed by anyone who wants to use it” for AI development, and that had he known OpenAI had trained on paywalled content, he would have required the company to retrain its models.

Microsoft Defends Its Position

A Microsoft spokesperson said Hecht’s comments “reflect one employee’s perspectives” and that the company’s position is set out in its court filings, which “explain why these transformative uses are consistent with copyright law and why Copilot is not a substitute for publishers’ journalism.”

Nadella’s remarks addressed changes in how people consume information and are “perfectly consistent” with the company’s legal stance, the spokesperson added, saying those observations “should not be confused with conclusions about copyright questions before the Court”.

In its own brief, Microsoft argued that the use of published content to train large language models was significantly transformative. “Copyright law does not permit rightsholders to block transformative technologies like LLMs; it encourages such uses on the expectation that rightsholders will adapt and the public will be better off for it,” the company stated.

Representatives for OpenAI and Ziff Davis did not immediately respond to requests for comment.

Source
Image: cnet.com

The UK tech briefing

Smartphones, AI, computing and deals — the essential stories without the noise.

Mailing provider can be connected when your UK list is ready.

Shop on Amazon UK — Discover deals Shop on Amazon UK — Discover deals