Majorregulation policyMicrosoft

Microsoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings reveal

Published
Sep 17, 2026 19:46 UTC
Also in this story:The New York TimesOpenAI

Brent Hecht, Microsoft’s director of Applied Science, characterized AI scraping as "the largest theft of labor in human history" in a January 2023 internal memo. This statement emerged in unredacted filings related to a copyright lawsuit involving The New York Times, which claims that OpenAI's models were trained on 91,692 copies of its works, alongside content from other publishers. The lawsuit highlights that 2 million documents from nytimes.com were included in a dataset derived from Common Crawl, a free web data repository.

In the wake of Microsoft's Copilot, The New York Times reported a 93% drop in click-through rates, indicating a significant impact on traffic and revenue. Satya Nadella, Microsoft CEO, emphasized that any paywalled content should be licensed for use in AI training, stating, "anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training." This follows Nadella's earlier testimony regarding the licensing of such content.

Nick Turley, head of ChatGPT at OpenAI, warned of an "existential threat" posed by generative AI products, while Greg Brockman, OpenAI's president, acknowledged that their models are "excellent at news." Hecht's internal presentation in January 2024 discussed the decline in performance attributed to AI products, reinforcing concerns about the potential disruption to employment in the news sector.

The implications of these developments are significant for AI practitioners, as they may need to reassess the datasets they use for training models, particularly those that include copyrighted material. This situation follows ongoing discussions about AI regulation and the ethical use of data, as highlighted in previous Turing Wire coverage.

Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.

Source: TechCrunch AI