AI

Microsoft’s Own Files Called AI Scraping the ‘Largest Theft of Labor in Human History’

Adrian Kessler
Add us on Google

A Microsoft scientist put a name to what the company was building, and the name was not flattering. Reviewing how large language models learn from the open web, he called it “the largest theft of labor in human history.” The line was written for internal eyes. It is now a court exhibit.

The phrase is landing as a gotcha: the AI giant’s own employee calling the business model theft. That reading is too small. What the unsealed filings actually show is a company that understood the mechanism it was monetizing, described that mechanism in plain language, and shipped the product anyway. The interesting part is not the insult. It is the diagram underneath it.

The document belongs to Brent Hecht, Microsoft’s director of applied science, and it surfaced in the copyright case The New York Times brought against OpenAI and Microsoft. Hecht argued that training models on published work without payment amounted to “an astonishing theft of unprecedented proportions.” A separate internal presentation went further into the machinery: it warned that Microsoft’s AI content strategy had started a “doom loop” that would “hurt the performance of our models and the entire web at the same time.” Large language models, one line reads, are “a product that destroys its supply chain.”

That last sentence is the whole story in seven words. A chatbot that answers a question inside the search box removes the reason to click through to the page it learned from. Fewer clicks mean less revenue for the sites that produce the writing, which means fewer of them survive, which means less fresh human text for the next model to train on. Microsoft did not theorize this loop from the outside. Its own data measured it: Copilot’s answer engine drove click-through rates to the Times’ site down by as much as 93 percent against ordinary search.

The filings also quantify the input side of the trade. OpenAI’s mid-training data held more than 91,000 copies of works from the Times, the Daily News, and the Center for Investigative Reporting. One dataset drawn from Common Crawl carried over two million documents from nytimes.com alone. A collection labeled Project Mango contained at least 160,000 unique works from news publishers. The filings describe paywalls bypassed and copyright notices stripped from the text before it was fed in.

OpenAI’s own executives were not naive about the effect. Nick Turley, who runs ChatGPT, is quoted describing products that are “largely substitutive” for journalism and calling the shift an “existential threat” to publishers. It is the same conclusion Hecht reached, filed under a different job title.

Here is why the language matters beyond embarrassment. Microsoft has distanced itself from the remarks, telling the court that its actual position complies with copyright law, and neither company would comment on the leak. But a fair-use defense turns partly on what a company knew and intended. Documents in which its own scientists name the harm — theft, a supply chain destroyed, a web hollowed out — are not opinions the company can simply disown. They read as evidence that the mechanism was understood before it scaled.

The web these models learned from was built by people who assumed a reader on the other end. Microsoft’s files show it knew, in writing, that the reader was being engineered out, and kept building.

Tags: , , , , ,

Add us on Google

Discussion

There are 0 comments.