Core Event: Secret Documents Reveal Executives’ AI News Training Warnings
During ongoing copyright litigation brought by The New York Times and other news outlets against Microsoft and OpenAI, a batch of previously classified internal documents has been formally disclosed. These files reveal deep internal concerns about using news content for AI training, with key facts as follows:
- Brent Hecht, Microsoft’s Director of Applied Sciences, warned in multiple documents that training AI on news content constitutes “unprecedented large-scale theft,” possibly “the largest labor theft in human history”
- Hecht questioned the legality of Microsoft and OpenAI’s “Fair Use” defense, arguing bulk scraping contradicts the principle of fair use
- Microsoft internally proposed a “doom loop” risk: AI products erode news site traffic → outlets cut content production → AI models lose reliable information sources
- Microsoft’s internal data shows: click-through rates for some of the involved news outlets dropped by 83% to 93%, with others falling 51% to 94%
- Nick Turley, OpenAI’s ChatGPT lead, acknowledged that commercial AI products replacing publishers pose a “existential threat” to media organizations
Critical Evidence: Traffic Decline and Internal Doubts
Microsoft’s internal data has partially validated the “doom loop” theory in practice. CEO Satya Nadella confessed in testimony that AI chatbots divert user attention, effectively “taking clicks away from news websites.” This trend manifests in concrete figures: some of the involved news outlets experienced over 80% click loss, with worst cases reaching 94% decline.
A second key revelation concerns content reproduction. Plaintiffs including The New York Times reported that testing Microsoft and OpenAI’s AI systems repeatedly yielded substantial verbatim reproduction of news articles. This raises questions about the effectiveness of training filters.
Hecht’s internal documents disclosed Microsoft developed a filtering mechanism that could make it “harder for copyright holders to understand what content was used in training.” This detail aligns with plaintiffs’ accusation that Microsoft “did not take adequate measures to prevent infringing outputs.”
Diverging Stances: Legal Defense vs. Internal Awareness
The most striking contradiction lies between internal warnings and public defenses. Microsoft’s spokesperson maintains its AI products qualify as transformative fair use that “do not replace news websites,” and insists Hecht’s documents reflect “only one employee’s personal views, not legal analysis or company policy.”
Yet internal records show executives like Hecht and Turley long grasped the severity. Turley internally described ChatGPT as “basically substitutive,” predicting this effect would intensify with capability gains. Regarding user behavior, his assessment was definitive: users have “no sufficient reason” to click original links.
The following table summarizes key witnesses and their internal positions:
| Person | Role | Internal Document Position |
|---|---|---|
| Brent Hecht | Microsoft Applied Sciences Head | News data training = “large-scale theft”; Fair Use defense questionable; Filtering obscurates copyright transparency |
| Satya Nadella | Microsoft CEO | Acknowledged chatbots reduce news site visits |
| Nick Turley | OpenAI ChatGPT Lead | AI replacement of publishers = “existential threat”; Product is “basically substitutive” |
Recommendations for Stakeholders
For news organizations: Prioritize testing AI output originality while exploring cooperative models beyond litigation. Some publishers now embed digital watermarks to identify processed content.
For AI end users: Recognize information drawn from AI may originate from contested data sources; verify critical information against primary sources rather than relying solely on AI platforms.
Final Thoughts
This document disclosure exposes fundamental tensions between rapid AI advancement and established intellectual property norms. As technological capability outpaces legal frameworks, rebuilding data-use transparency and fair value allocation must precede broader restoration of trust.
(Word count: 498)
