Microsoft exec labeled AI scraping 'the largest theft of labor in human history'
Original: Microsoft exec called AI scraping 'the largest theft of labor in human history'
Why This Matters
Internal admissions of harm directly undercut AI companies' fair use defense and could reshape copyright litigation outcomes.
Unredacted filings in the NYT v. OpenAI/Microsoft lawsuit reveal that a senior Microsoft executive privately called AI training scraping 'theft,' while OpenAI's own leadership acknowledged its models posed an 'existential threat' to publishers. Microsoft's Copilot was found to cut NYT click-throughs by up to 93%.
New unredacted court filings in The New York Times' three-year-old copyright lawsuit against OpenAI and Microsoft contain striking internal admissions. A top Microsoft executive privately described the companies' AI training practices as 'the largest theft of labor in human history.' OpenAI leadership separately acknowledged that its models posed an 'existential threat' to the publishers and journalists whose content trained them.
The filings also allege that the companies bypassed paywalls undetected, built training datasets through mass scraping, and deliberately stripped copyright notices from training data. Microsoft's own internal data shows its Copilot 'answer engine' caused click-through rates to NYT's domain to drop as much as 93% compared to traditional Bing search results.
A January 2024 internal presentation by Microsoft's director of Applied Science, Brent Hecht, called this a 'doom loop' that would 'hurt the performance of our models and the entire web at the same time.' The document noted: 'It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business.'
Microsoft CEO Satya Nadella testified in deposition that 'anything that is paywalled should be licensed by anyone who wants to use it.' It's worth noting that the new disclosures come from The Times' own legal brief, not from underlying sealed exhibits, meaning the quotes currently lack full original context. Judges have so far been broadly favorable to AI companies' 'fair use' arguments, and the Trump administration recently filed a brief in OpenAI's defense.