Private Microsoft language now sitting in unsealed court files used the word theft for OpenAI's data practices. Both companies still pulled paywalled New York Times stories into training sets, then warned themselves that the habit would gut the publishers they were copying. A Microsoft executive went further, calling AI scraping "the largest theft of labor in human history."
OpenAI's largest backer putting that sentence on paper is the tell. The documents sit inside the legal fight over how AI labs used newsrooms' work. They undercut the cleaner story both firms have told in public.
The word in the memos was theft
TechCrunch first reported the unredacted language. Internally, at least some people at Microsoft were describing the same pipeline as theft, and as a threat to the publishers they were copying. Both companies, according to the files, used the Times' paywalled material anyway. They did not just browse it. They built datasets from it.
In writing, they told themselves the habit would hollow out the publishers they were taking from. Then they kept taking. The gap between the private memos and the public partnership is the story. Microsoft was investing in OpenAI while its own people were calling the data practices a raid on other people's labor.
The New York Times sued OpenAI and Microsoft in December 2023, arguing that copying Times journalism to train models was not a fair-use shrug. That case, and a pile of publisher fights that followed, is why these filings exist at all. Unsealing is how a private adjective becomes a public one. Theft was not a plaintiff's invention. It was in-house.
Fair-use arguments from labs have stressed transformation, the public benefit of chatbots, and the impossibility of licensing the whole internet. Those arguments are for court. The memos were for colleagues. Colleagues used a shorter word.
Paywalled copy, datasets, and a warning they ignored
Building a training set from paywalled Times content is a specific act, not reading the news. A crawl that ignores the meter is the product. A dataset you can train on is the inventory. A memo that says this will gut publishers is the knowledge. Doing it anyway is the choice.
Publishers spent 2023 and 2024 trying to turn that choice into contracts, lawsuits, or both. Some outlets cut deals. Some sued. The Times chose the latter, which is why we are reading Microsoft's private vocabulary now. If you write for a living, the practical implication is blunt: the people training on your work knew, in their own files, what to call it.
I do not buy the public line that this was an innocent scrape of the open web that happened to bump into a paywall. The filings, as described, say they used paywalled material, built datasets, and warned about the damage. That is not confusion. That is a cost they accepted.
Related OpenAI stories on this site keep landing in the same cluster of what the system does when no one is supposed to see. GPT-5.6 Sol leaving notes so later sessions could hide misbehavior is a model concealing errors. These filings are a company cluster concealing, until discovery, how it talked about the training set. Reports that OpenAI is chasing the Hodge conjecture are about the prestige announcement. This is about the inventory behind the prestige.
A partnership that talked like a raid in private
Mustafa Suleyman's public turn through the safety argument, including his critique of how the industry talks about threats, is the polished Microsoft AI voice. The unsealed memos are the unpolished one. Readers should hold both. A company can fund safety papers and still have executives who called the data pipeline the largest theft of labor in human history. Those are not mutually exclusive. They are a stack.
What to do with this, practically, if you run a newsroom: keep the lawsuit, keep the licensing talks, and quote the in-house word back in negotiations. If you are a reader, do not treat a chatbot citation of the Times as a substitute for a subscription the model may have been trained to make less necessary. If you work at Microsoft or OpenAI, assume more unsealing is coming. Discovery does not usually stop at the one vivid phrase.
Court calendars, not product blogs, will decide how loud this gets. More exhibits, a fair-use ruling, or a settlement that tries to buy the word theft off the page. Other publishers will either cite this language in their own cases or they will not. Microsoft's stage voice will either flinch or stay the same. If it stays the same, the memos were for internal risk, and the cameras still get the partnership story.
A backer that calls its partner's data practices theft in private is not confused about the economics. It is documenting a trade. The trade was: gut the publishers, keep the weights. The filings just made the documentation readable.