// ARS TECHNICA — INTELLIGENZA ARTIFICIALE
Microsoft exec called AI scraping the “largest theft of labor in human history”
Microsoft, OpenAI emails reveal fear of AI “doom loop” killing news orgs.
For years, Microsoft and OpenAI have fought to keep certain information out of the public eye in their fight with news organizations that have accused the AI firms of teaming up to violate copyright laws by stealing tons of news content to train AI.
However, now the details that should never have been marked confidential are starting to leak. In a motion for summary judgment that was unsealed Thursday from news plaintiffs led by The New York Times, internal documents are exposed that news groups alleged show exactly how Microsoft and OpenAI viewed the threat to news before unleashing new AI products like ChatGPT and Copilot.
Perhaps most explosively, Microsoft Director of Applied Science Brent Hecht repeatedly warned in documents that scraping news for AI training was “an astonishing theft of unprecedented proportions,” calling it perhaps the “largest theft of labor in human history,” news orgs said. In another document, Hecht contradicted Microsoft and OpenAI’s argument that training AI on news content is fair use, suggesting that the plan to widely scrape news made “a complete mockery of the idea of ‘fair use.’”
Over at OpenAI, ChatGPT head Nick Turley wrote in an internal message that publishers would face an “existential threat” from commercial products trained on news content that can be used to substitute news providers. One Microsoft document even described a “doom loop,” news orgs said, “that will hurt the performance of our models and the entire web at the same time.”
“It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain,’” that document said.
Data from both firms shows that this prediction was accurate. Microsoft recorded 83–93 percent drops in click-through rates for some news plaintiffs, and 51–94 percent drops for others. Add to that reporting on low click-through rates from ChatGPT search results and news organizations’ own reporting on traffic declines. Suddenly, it becomes easier to see how declining news revenue could ultimately rob chatbots of the abundant streams of reliable information that supposedly makes them such groundbreaking tools.
Meanwhile, “almost no one intended for content they created to be used in this fashion, nor are they compensated for its use,” Hecht acknowledged in a Microsoft document.
News organizations say they’re ready to go to trial because there’s so much “compelling evidence of substitution.” If they can prove that chatbots are replacing them in their own markets, while serving to spit out excerpts of articles verbatim, they think that one-two punch may eviscerate Microsoft and OpenAI’s fair use arguments.
“The future not just of journalism but of responsible AI too depends on preserving incentives for humans to produce the creative works on which a healthy society depends,” news groups argued.