Microsoft Exec Called AI Scraping 'the Largest Theft of Labor in Human History,' Unredacted NYT v. OpenAI Filings Show
Newly unredacted material in The New York Times v. OpenAI and Microsoft copyright lawsuit includes internal Microsoft and OpenAI communications that undercut the companies' fair-use defense. TechCrunch reports the filing says a Microsoft director of Applied Science, Brent Hecht, described the companies' training-data acquisition in a January 2023 memo as "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history."
What the filings allege
- The companies allegedly bypassed paywalls undetected, built training datasets via mass scraping, and stripped copyright notices from training data.
- OpenAI's mid-training datasets alone contain more than 91,692 copies of works published by the NYT, Daily News, and Center for Investigative Reporting.
- A Common Crawl-derived dataset included more than 2 million documents from nytimes.com alone.
- Microsoft's own data shows its Copilot "answer engine" dropped click-through rates for the NYT domain by as much as 93% versus traditional Bing search.
The 'doom loop' internal doc
A January 2024 Microsoft presentation by Hecht described the traffic decline as a "doom loop" that would "hurt the performance of our models and the entire web at the same time," and warned: "It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain.'" A separate Microsoft document flags a "real risk" that generative AI could "significantly disrupt the employment of the very people who generated the data on which the foundation model was trained."
Statements that cut against fair use
- Microsoft CEO Satya Nadella testified this year that "anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training," and said he would have required OpenAI to retrain if he knew it scraped paywalled data.
- OpenAI's head of ChatGPT, Nick Turley, wrote internally that publishers face an "existential threat" because the chatbot is "largely substitutive" and "will get more and more substitutive as they get better."
- OpenAI President Greg Brockman described the models as "excellent at news." Nadella agreed under oath that chatting with a bot substitutes for visiting the source website.
Context
The case has run three years; judges have generally leaned toward AI companies on fair use, and the Trump administration filed a brief this month defending OpenAI's unlicensed training. Caveat: much of the new material comes from The Times' own brief, not the underlying sealed exhibits, and the quoted passages are presented without their original context. For developers shipping RAG or training pipelines, the discovery here is a reminder that data provenance and licensing terms are becoming courtroom evidence, not just engineering hygiene.
📖 Read the full source: HN AI Agents
👀 See Also

OpenClaw Client Adds Cost Tracking and Per-Agent Spending Limits
New release adds spending caps per agent, live usage UI with circular progress bar, sub-agent management, skill toggling, and per-agent model selection.

OpenAI Frontier Models and Codex Now Available on AWS
OpenAI's frontier models and Codex are now generally available on AWS, letting enterprises use OpenAI via their existing AWS environments and procurement workflows.

SDNY Court Rules AI-Generated Legal Documents Not Protected by Privilege
Judge Jed S. Rakoff ruled that 31 documents generated using Anthropic's Claude AI tool were not protected by attorney-client privilege or work product doctrine, marking the first such court decision on AI-generated legal materials.

UK Home Office Used AI-Hallucinated Document to Deny Asylum Claim, Judge Finds
A senior UK judge ruled the Home Office likely used AI-generated hallucinated evidence—a non-existent country policy note—to refuse an asylum claim. The missing document was never found, and the judge compared the practice to relying on bogus evidence.