Microsoft exec called AI scraping 'the largest theft of labor in human history'

New unredacted legal filings reveal how Big AI really thinks about the content it ingested.

Published: 2026-09-18 Category: Quick Take Sources: TechCrunch

The admission that was always there

Three years into The New York Times' copyright lawsuit against OpenAI and Microsoft, unredacted material has finally surfaced — and it contains the kind of language the fair-use defense has spent years trying to keep out of the record. A top Microsoft executive privately described the companies' training practices as "theft," and OpenAI's own leadership said its models posed an "existential threat" to the publishers whose work trained them.

The unsealed detail is a gift to plaintiffs: it hands them the companies' own words, in their own internal documents, admitting the thing their public posture has long denied. That the quotes come from The Times' brief rather than underlying exhibits matters — they're presented without original context — but they're on the record now. The genie isn't going back into the bottle.

The numbers are the story

The scale of copying is what makes "fair use" so hard to swallow. OpenAI's mid-training datasets alone contained more than 91,692 copies of works from the NYT, Daily News, and the Center for Investigative Reporting. A Common Crawl-derived dataset included more than 2 million documents from nytimes.com alone. The filings allege the content was obtained by bypassing paywalls, mass scraping, and deliberately stripping copyright notices from training data.

A January 2023 internal memo from Microsoft's Brent Hecht called it "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history." Internal Microsoft data showed its Copilot "answer engine" cut click-through rates to The New York Times by as much as 93% versus traditional Bing — a "doom loop" that would hurt both the models and the web simultaneously.

The fair-use contradiction

The heart of the fair-use test asks whether use substitutes for and harms the market for the original work. OpenAI's head of ChatGPT, Nick Turley, wrote internally that publisher-facing chatbots are "largely substitutive" and will "get more and more substitutive as they get better." CEO Satya Nadella testified that "anything that is paywalled should be licensed," and said he'd have forced OpenAI to retrain had he known paywalled content was scraped. That language — products that compete with, rather than transform, the original — cuts hard against the fair-use shield.

The irony is that judges have largely leaned toward AI companies so far, and the Trump administration filed a brief this month backing OpenAI's unlicensed training. The battle over whether training on copyrighted work is "fair use" or "theft" is the defining legal fight of the generative-AI era, and these filings put the companies' own vocabulary squarely on the scales.


Analysis based on reporting by TechCrunch. Read the full story at techcrunch.com.