Microsoft Says 'Virtually Nobody' Was Grabbing NYT Articles Through Copilot - and That's Its Fair-Use Argument
In the landmark copyright fight, Microsoft is leaning on chat logs to argue its AI doesn't reproduce enough copyrighted text to matter.
Published: 2026-09-05 Category: Quick Take Sources: The Verge
The Story
As part of discovery in the consolidated copyright lawsuit against Microsoft and OpenAI, Microsoft provided 8.2 million Copilot chat logs to an expert hired by the news publishers. The logs, Microsoft claims, were specifically chosen "because they hit on keywords implicating use of News Plaintiffs' websites" — meaning they were the most likely to contain the publishers' works.
The resulting analysis, Microsoft says, shows that only 59,545 of those logs contained at least 16 words in common with news content used to ground the AI model. An expert for the Center for Investigative Reporting found 51 instances of "substantial overlap" with CIR work. In the authors' suit, the 8.2 million conversations had just 24 responses containing at least 30 matching words, and only 10 of the 212 books evaluated had any matches.
The Argument
Microsoft's numbers are meant to bolster a fair-use defense: while systems like Copilot rely on copyrighted material, the resulting systems are used for significantly different purposes than the original. The fact that they sometimes reproduce sections of text, Microsoft concludes, "hardly undermines the transformative purpose of LLM training."
The New York Times disagrees sharply. "The documents and testimony uncovered during discovery lead to only one conclusion: Microsoft and OpenAI stole from The New York Times to make commercial products that substitute for its journalism, threaten its business, and undermine its industry," said lead counsel Ian Crosby. "We look forward to Microsoft and OpenAI being held accountable for their theft."
Why It Matters
This is a pivotal moment in the AI copyright wars. Microsoft filed its motion for summary judgment on Friday, arguing the case should end at an early stage. The Trump administration also filed a statement of interest in the NYT case this week — supporting OpenAI. If the judge sides with Microsoft and OpenAI, the case ends; if not, the legal battle continues in court.
The framing is clever: Microsoft is trying to shift the question from "did you train on copyrighted works?" (which is undisputed) to "did your product actually reproduce them in a way that harms the market?" That's a much friendlier battlefield for the AI companies.
The Takeaway
The publishers' counterargument is that even a small number of reproductions matters when the entire business model depends on training on their work without permission. The numbers Microsoft cites — 59,545 logs, 24 responses, 10 books — may sound small, but the plaintiffs will argue they're the tip of an iceberg that only exists because of wholesale copying. This case, more than any other, will define whether training on copyrighted content is fair use or theft.
Source: The Verge, "Microsoft says virtually nobody was grabbing NYT articles through its chatbot" (September 4, 2026).