Microsoft Exec Called AI Scraping 'the Largest Theft of Labor in Human History,' Unredacted NYT v. OpenAI Filings Show

✍️ OpenClawRadar📅 Published: September 18, 2026🔗 Source
Ad

Newly unredacted material in The New York Times v. OpenAI and Microsoft copyright lawsuit includes internal Microsoft and OpenAI communications that undercut the companies' fair-use defense. TechCrunch reports the filing says a Microsoft director of Applied Science, Brent Hecht, described the companies' training-data acquisition in a January 2023 memo as "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history."

What the filings allege

  • The companies allegedly bypassed paywalls undetected, built training datasets via mass scraping, and stripped copyright notices from training data.
  • OpenAI's mid-training datasets alone contain more than 91,692 copies of works published by the NYT, Daily News, and Center for Investigative Reporting.
  • A Common Crawl-derived dataset included more than 2 million documents from nytimes.com alone.
  • Microsoft's own data shows its Copilot "answer engine" dropped click-through rates for the NYT domain by as much as 93% versus traditional Bing search.

The 'doom loop' internal doc

A January 2024 Microsoft presentation by Hecht described the traffic decline as a "doom loop" that would "hurt the performance of our models and the entire web at the same time," and warned: "It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain.'" A separate Microsoft document flags a "real risk" that generative AI could "significantly disrupt the employment of the very people who generated the data on which the foundation model was trained."

Ad

Statements that cut against fair use

  • Microsoft CEO Satya Nadella testified this year that "anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training," and said he would have required OpenAI to retrain if he knew it scraped paywalled data.
  • OpenAI's head of ChatGPT, Nick Turley, wrote internally that publishers face an "existential threat" because the chatbot is "largely substitutive" and "will get more and more substitutive as they get better."
  • OpenAI President Greg Brockman described the models as "excellent at news." Nadella agreed under oath that chatting with a bot substitutes for visiting the source website.

Context

The case has run three years; judges have generally leaned toward AI companies on fair use, and the Trump administration filed a brief this month defending OpenAI's unlicensed training. Caveat: much of the new material comes from The Times' own brief, not the underlying sealed exhibits, and the quoted passages are presented without their original context. For developers shipping RAG or training pipelines, the discovery here is a reminder that data provenance and licensing terms are becoming courtroom evidence, not just engineering hygiene.

📖 Read the full source: HN AI Agents

Ad

👀 See Also