In a January 2023 internal memo, Microsoft’s director of applied science Brent Hecht called the scraping of the open web to train language models “the largest theft of labor in human history.” The line surfaced in a New York Times filing unsealed on September 17, in the paper’s copyright suit against OpenAI and Microsoft.
Key Takeaways
- Hecht also described the practice as “an astonishing theft of unprecedented proportions.”
- Internal data showed Copilot cut click-through to the New York Times by as much as 93% compared with classic Bing search.
- OpenAI’s mid-training datasets held more than 91,692 copies of works from the Times, the Daily News and the Center for Investigative Reporting.
Have an AI Sum Up This Article
ChatGPTThe Microsoft memo that just left the sealed file
The Times sued in 2023 in federal court in the Southern District of New York. It alleges that OpenAI and Microsoft copied millions of its articles without permission to train commercial systems built to compete with news outlets.
What changed this week is a previously redacted motion from the paper, now readable in the version filed with the court and posted on CourtListener. The Times’ lawyers use internal documents from both companies to attack their fair use defense.
US courts weigh fair use case by case, looking in particular at whether the use is transformative and at its effect on the market for the original work. The internal documents the Times quotes are aimed squarely at that second factor.
The most quoted exhibit is Hecht’s January 2023 memo, written while he ran applied science at Microsoft. He called industry-wide scraping of the web for large language models “an astonishing theft of unprecedented proportions” and “the largest theft of labor in human history.”
He reportedly went further, writing that allowing such copying “would make a complete mockery of the idea of ‘fair use.'” For the Times, that is the whole point: the words come from inside Microsoft and they target the exact legal argument the company now uses in court.
The filing also puts a number on what Copilot did to the paper’s traffic. Internal data showed click-through rates to the New York Times dropping as much as 93% compared with traditional Bing results.
A Microsoft presentation described the effect as a “doom loop” that would “hurt the performance of our models and the entire web.” Put plainly, the company expected an assistant that answers in place of websites to eventually starve the content it feeds on.
That loop already shows up in the data. Over the summer AI-written text reached one in three pages on the web, so publishers get fewer clicks while machines produce more of what gets indexed.
What the exhibits say about OpenAI and Satya Nadella
The file goes well beyond Microsoft. OpenAI’s mid-training datasets contained more than 91,692 copies of works published by the New York Times, the Daily News and the Center for Investigative Reporting.
A dataset derived from Common Crawl, the giant public archive of the open web, held more than 2 million documents from nytimes.com alone. The Times also alleges that employees circumvented paywalls and stripped copyright notices from the data before training. The Daily News and the Center for Investigative Reporting sit alongside the Times in those figures, which widens the dispute beyond a single newsroom.
OpenAI’s internal messages point the same way. ChatGPT head Nick Turley wrote that publishers face an “existential threat” from chatbots. President Greg Brockman described the models as “excellent at news” and said the products “will get more and more substitutive” as they improve.
That word carries weight in a copyright case. The more a product substitutes for the work it copied, the harder it is to argue the use was transformative, one of the pillars of any fair use defense.
Satya Nadella’s testimony rounds it out. Microsoft’s CEO said paywalled content “should be licensed by anyone who wants to use it” for grounding or training, and that he would have required OpenAI to retrain its models had he known paywalled material was scraped. Coming from the chief executive of a defendant, that position narrows the room for the fair use argument his own lawyers are making.
Grounding means feeding documents to a model at the moment it answers, for instance when an assistant condenses a story published that morning. A license that only covers training would leave all of those real-time uses unaddressed.
Neither OpenAI nor Microsoft commented to reporters. A judge still has to decide whether the case goes to trial, and the lawsuit joins a crowded docket for OpenAI after the Elon Musk trial, where trust sat at the center of the verdict.
More articles on Horizon
- OpenAI Models Tried to Hide Their Mistakes
- DeepMind Institute Tackles the Rise of Human-Level AI
- Review a Contract With Claude Without Leaking Data
Why the Microsoft quotes shift the fight with publishers
For people using Copilot or ChatGPT, nothing changes today. The services stay online, and unsealed exhibits are not a ruling. They describe internal practices and opinions whose legal weight still has to be tested in court.
The impact lands in licensing talks between labs and news outlets. Every internal quote that admits substitution strengthens the hand of publishers asking for paid licenses rather than a link at the bottom of an answer.
Nadella’s deposition hands other plaintiffs an argument too. If the head of Microsoft says paywalled content needs a license for grounding as well as training, any assistant summarizing paywalled articles will have to explain how it gets access.
Teams plugging AI assistants into their research and monitoring tools now have a contract question to ask. Where the summarized content comes from, and under which license it is served, belongs on the same checklist as data privacy.
On the competitive side, Microsoft and OpenAI carry these exhibits alone, even though every lab trained on the open web. Model makers that signed licensing deals with news organizations can now pitch them as legal cover as much as a business arrangement.
Publishers walk away with a number they never had from a Big Tech source: a click-through drop of up to 93%, measured by Microsoft itself. It covers one publisher and compares Copilot with classic Bing, yet it gives negotiators an order of magnitude.
Traceability is also moving on the lab side, as shown by Anthropic opening its Claude text detection tool to media outlets. Provenance is slowly turning into a product feature that publishers can point to when they negotiate.
The two companies have yet to answer these excerpts in public, and until a judge rules on fair use they can still frame the memo as one employee’s view in an internal debate. Hecht’s phrase will be hard to walk back all the same, three years after he wrote it. What the court makes of these exhibits will shape the price of news for every assistant that answers questions about current events.
Follow the story on Horizon.


