Core Event
Sealed court documents from The New York Times’ lawsuit against OpenAI and Microsoft have been recently unsealed, revealing alarming internal admissions. The materials show that both companies were aware their AI training practices posed existential threats to publishers yet proceeded anyway driven by commercial ambitions.
Key hard facts:
- Documents originate from The New York Times’ copyright infringement litigation
- Content reflects internal discussions from approximately 2022–2023
- Key figures cited include OpenAI co-founder Greg Brockman, CEO Sam Altman, and Microsoft CEO Satya Nadella
- Internaladers expressed legal concerns about training data harvesting practices
- The case remains pending with no final judgment issued
Internal Warnings: ‘Doom Loop’ and ‘Labor Theft’

Most striking are the statements by Microsoft’s Director of Applied Science, Brent Hecht. He described ChatGPT and Copilot’s data harvesting as the “largest theft of labor in human history” and stated Microsoft’s legal defense makes a “complete mockery of the idea of ‘fair use’.” Microsoft has since distanced itself, claiming Hecht’s views are “one employee’s individual perspective” not company policy.
An internal Microsoft document explicitly uses the term “doom loop,” warning that the AI content strategy has “started a ‘doom loop’ that will hurt the performance of our models and the entire web at the same time,” noting it is “highly unusual that an end-product threatens the economic foundations of its essential suppliers.”
OpenAI’s own documents reveal contradictions: internal memos stress that “prevention of memorization” is critical to “minimize copyright violations,” yet employees acknowledged GPT-4 memorized a ton of data and therefore will be insanely good at regurgitation. Specific examples cited include verbatim reproduction of articles from The New York Times, The Mercury News, The Denver Post, LifeHacker, and Eurogamer.
Profit Motives and Tracking Failures
Despite public statements supporting licensing, internal docs show otherwise. While Nadella publicly stated “anything that is paywalled should be licensed,” an OpenAI representative admitted being “unaware of any effort to detection or remove paywalled content from training data”—a stark gap between rhetoric and practice.
OpenAI Policy Director Jack Clark warned the company was “creating systems that substitute for the labor of the people that define the ‘culture’ of society,” internally labeling ChatGPT “the modern newsstand.” Nick Turley, OpenAI’s ChatGPT lead, stated users obtaining AI answers had “no good reason to click” on source links, causing publishers to suffer referral traffic drops of up to 60% according to OpenAI’s own internal analysis.
Microsoft’s spokesperson cautiously notes Nadella’s testimony reflects “broad principles” unrelated to legal copyright arguments, yet the documents show clear foresight: LLMs are “a product that destroys its own supply chain” because the AI output itself substitutes for the original training data.
Technical and Legal Tensions

The case highlights a fundamental copyright tension in the AI era: large language models require massive text datasets, much of it copyright-protected. Under U.S. fair use doctrine—which weighs purpose, nature, amount used, and market impact—the documents suggest internal confidence was low, yet commercial potential proved decisive.
Greg Brockman’s focus appears evident: his concern was “gazillions of dollars” of potential revenue, not legal risk—suggesting commercial incentives overrode compliance.
Reader Recommendations

- Content creators and publishers: These filings offer strong evidentiary support for potential copyright claims. Maintain systematic timestamped records of original content for future enforcement.
- Enterprise AI adopters: Internal documents confirm verbatim reproduction remains a high-risk issue; professional-domain outputs (legal, medical, financial) require human verification.
- General users: Even improved models cannot fully eliminate data memorization; treat AI-generated summaries as starting points, not authoritative sources.
Final Note
When internal documents label a product the “destroyer of its own supply chain,” the cost of technological acceleration becomes unequivocal—efficiency leaps and systemic risk are two sides of the same coin. This case’s outcome may redefine fair use for the digital age, with ramifications extending far beyond these two litigants.

