A Microsoft Exec Called AI Scraping 'The Largest Theft of Labor in Human History'
Some quotes are hard to walk back once they're in a court filing. This week, legal briefs unsealed in the New York Times' copyright lawsuit against OpenAI and Microsoft revealed one that's going to follow both companies for a while: a Microsoft director, in internal communications, describing AI training data scraping as "the largest theft of labor in human history."
That's not a critic or a journalist saying that. That's someone inside Microsoft.
What's Actually in the Filings
The Times is asking a judge for summary judgment in its case against OpenAI and Microsoft, and the legal brief leans heavily on internal documents from both companies to make its case. A few of the details are striking:
- The Microsoft director's "largest theft of labor" comment was made internally, not for public consumption — which is part of why it's landing so hard now that it's on the record.
- OpenAI's own head of ChatGPT reportedly described the New York Times as an "existential threat to publishers," an odd thing to say about the party your company is being sued by, but revealing about how OpenAI's leadership talks about the news industry internally.
- The filing claims Microsoft's Copilot has cut the Times' referral traffic by as much as 93% compared to what it would get from a plain Bing search — a huge number if it holds up, since it suggests AI answers are replacing clicks to the original source almost entirely.
- Satya Nadella is quoted from testimony suggesting that content behind paywalls should require licensing — a position that, if taken at face value, cuts against OpenAI's "fair use" defense in the same case.
Why This Matters Beyond One Lawsuit
Copyright suits against AI companies aren't new — publishers, authors, and artists have been filing them since 2023. What makes this one different is that the plaintiffs aren't just alleging harm from the outside. They're quoting the defendants' own people describing the harm from the inside.
That's a much harder thing for a jury — or a judge weighing summary judgment — to wave away. "Fair use" arguments tend to hinge on whether the copying was transformative and whether it damaged the market for the original work. A 93% traffic drop and an internal admission that content was taken without proper compensation both cut directly against that defense.
It also puts a number on something publishers have argued qualitatively for years: that AI chat interfaces don't just summarize the news, they replace the reason to visit the source at all.
The Bigger Tension Nobody's Resolved
Here's the uncomfortable part: the same companies being accused of mass-scraping content are also the ones selling the tools that make scraping, summarizing, and repackaging content trivially easy for everyone else. Microsoft and OpenAI aren't just defendants in a labor-and-copyright dispute — they're also the infrastructure the entire AI industry, including agent builders, runs on top of.
That tension doesn't go away because a judge rules one way or another. It's a live question for anyone building AI products right now: where did the training data come from, and who — if anyone — got paid for it?
What This Means if You Use OpenClaw
This story is a reminder that the AI tools layer and the AI data layer are two very different things, and it's worth knowing which one you're relying on.
OpenClaw isn't a foundation model company scraping the web to train something new — it's an open-source agent framework that lets you point existing models at your tools, your accounts, and your workflows. The value isn't in a training corpus someone else assembled; it's in what your agent actually does for you, task by task, with data you control.
That distinction matters more as lawsuits like this one play out. The safest place to be, as a builder or a user, is one where you know exactly what your agent is touching and why — not wondering what went into a black-box model you're renting.
The Bigger Picture
Whatever a court eventually decides about fair use and labor, this case is forcing something useful into the open: real numbers, real internal admissions, and a clearer picture of who actually benefits when AI repackages other people's work.
If that makes you want more control over what your own agents are doing and where their outputs come from, that's exactly the kind of transparency OpenClaw is built around.