Microsoft exec called AI scraping 'the largest theft of labor in human history,' new unredacted filings reveal
The quote is real, and it is Microsoft's own scientist saying it about its own biggest partner. What it proves legally is a separate, much harder question that a dramatic sentence cannot answer by itself.

The full argument.
Newly unredacted filings in The New York Times' copyright suit against OpenAI and Microsoft, made public September 17, 2026, quote Brent Hecht, Microsoft's director of Applied Science, calling AI training practices "an astonishing theft of unprecedented proportions" and, separately, "the largest theft of labor in human history." Both lines come from a January 2023 internal memo. A January 2024 presentation attributed to Hecht warns of a "doom loop" that would "hurt the performance of our models and the entire web at the same time," and adds: "It is highly unusual that an end-product threatens the economic foundations of its essential suppliers."
That is the sentence doing all the work in every headline today, including the one on this piece, because Microsoft is the company defending OpenAI's data practices in court while, on this evidence, one of its own scientists was privately calling those practices theft. That is a real, checkable tension. It is not, by itself, a verdict.
Here is what the filings, as reported, actually put on the record. OpenAI's mid-training datasets allegedly held more than 91,692 copies of content from the Times, the Daily News and the Center for Investigative Reporting. A Common Crawl-derived dataset allegedly contained upward of 2 million documents from nytimes.com. Microsoft is alleged to have shared data with OpenAI through internal projects code-named Taxi and Mango, with Mango said to contain more than 160,000 unique news articles. Researchers are alleged to have stripped copyright notices from training text so models would not reproduce them. In one exchange, an OpenAI employee reportedly told president Greg Brockman about a way to get around the Times' paywall; Brockman is said to have replied, "ah nice." And Microsoft's CEO, Satya Nadella, testified in a deposition that paywalled material "should be licensed by anyone who wants to use it," and that he would have forced OpenAI to retrain its models had he known paywalled content was scraped. Microsoft's own internal data reportedly showed Copilot's presence cutting click-through to Times content by as much as 93% versus plain Bing search — the number the Times needs to argue AI chat substitutes for, rather than merely draws on, the original reporting.
All of that is specific and, if accurately quoted, damning to read. None of it is something this outlet independently verified against the underlying exhibits, because the exhibits are still sealed. What is public is the Times' own legal brief characterizing those exhibits — the party with the largest possible incentive to select the most damning eleven words out of a hundred-page memo and put them in a filing headline. That does not make the quotes fake. Litigants quote real documents constantly. It means the selection itself is not neutral, and a reader should hold that thought at the same time as the quote.
Microsoft's official response, provided to TechCrunch, distances the company from Hecht's language without disputing that he wrote it: "Microsoft's position is set out in its court filings, which explain why these transformative uses are consistent with copyright law and why Copilot is not a substitute for publishers' journalism." Note what that statement does and doesn't do. It doesn't say Hecht is misquoted. It reasserts the legal argument — transformative use — that determines the case regardless of how anyone inside the building felt about it in a memo.
That's the actual news here, and it's worth separating from the theft framing. Fair use, as a legal doctrine, does not ask whether the acquisition process felt honest to the people doing it. It asks whether the resulting use is transformative and whether it harms the market for the original. Courts weighing that question in AI cases have so far leaned toward the AI companies, and the U.S. government filed a brief this month backing OpenAI's position in this case specifically. An employee's private disgust is evidence a jury might find persuasive on intent or bad faith — legal scholars quoted elsewhere note the admissions cut against the substitution argument OpenAI needs to win on fair use — but it is not itself the legal standard, and it is not proof the company's own defense will fail.
So: the quotes are real, sourced to a named Microsoft scientist, and worth taking seriously as evidence of what people inside these companies believed they were doing. The scale figures — the document counts, the click-through collapse — are the concrete, checkable part of the story, more useful to hold onto than the phrase built to go viral. What the filings prove about the actual copyright case is still, as of this writing, undecided, and no single memo settles it either way.