// THE VERGE — INTELLIGENZA ARTIFICIALE
Microsoft says virtually nobody was grabbing NYT articles through its chatbot
Posts from this topic will be added to your daily email digest and your homepage feed.
Posts from this topic will be added to your daily email digest and your homepage feed.
Posts from this topic will be added to your daily email digest and your homepage feed.
Fewer than 1 percent of more than 8 million chat logs regurgitated at least 16 words.
Fewer than 1 percent of more than 8 million chat logs regurgitated at least 16 words.
Posts from this author will be added to your daily email digest and your homepage feed.
Posts from this author will be added to your daily email digest and your homepage feed.
Microsoft’s Copilot rarely reproduces even full sentences from news articles and books, let alone substantive chunks that could substitute for the original, the company says in new legal filings as it fights copyright claims from publishers including The New York Times and book authors.
As part of the lawsuit’s discovery, Microsoft provided 8.2 million Copilot chat logs to an expert hired by news publishers. The logs, it claims, were specifically chosen “because they hit on keywords implicating use of News Plaintiffs’ websites, and therefore the most likely to contain News Plaintiffs’ works.” It says the resulting analysis shows that 59,545 of these contained at least 16 words in common with news content used to ground the AI model. An expert for the Center for Investigative Reporting found 51 instances of “substantial overlap” with CIR work in the dataset, Microsoft says. Similarly, an expert in the authors’ suit found that the 8.2 million conversations with Copilot only had 24 responses that contained at least 30 matching words. Only 10 of the 212 books evaluated had any matches, Microsoft claims. The Times, CIR, and Authors Guild did not immediately respond to requests for comment.
Microsoft argues that the numbers bolster its case that using copyrighted content for AI training datasets should be considered fair use. While systems like Copilot rely on using copyrighted material, it says, the resulting systems are used for significantly different purposes than the original. The fact that they sometimes reproduce sections of text, it concludes, “hardly undermines the transformative purpose of LLM training.”