Full Breakdown
Microsoft and OpenAI Face Internal Dissent Over News Scraping Amid Ongoing Copyright Lawsuit
By Drooid · · How we work
Core Event: Employees Label News Scraping “Largest Theft of Labor”
The same documents show OpenAI staff, including the head of ChatGPT, describing AI as an “existential threat” to publishers and predicting that AI products would become increasingly “substitutive.” Internal memos also describe a “doom loop” in which the models’ reliance on scraped content could degrade both the web and the models themselves.
Background & Context: Lawsuit and Industry Tensions
The New York Times filed a copyright lawsuit against Microsoft and OpenAI in late 2023, alleging that the companies scraped the newspaper’s articles without permission or payment to train AI systems. Eleven additional publishers have joined the suit, and the Seattle Times and Newsday filed parallel complaints in September.
Both companies argue that their use of the material qualifies as “fair use” because the AI-generated output is transformed and does not substitute for the original journalism.
Data & Statistics: Measurable Impact on Publisher Traffic
Court documents note an 83 % to 93 % decline in click-through rates for the New York Times and Daily News when users were served results from Microsoft’s AI-enhanced Bing search compared with the traditional Bing engine. The drop suggests that AI-driven search may be diverting traffic away from publisher sites.
Official Statements & Responses
In a deposition, Microsoft CEO Satya Nadella stated that any paywalled material should be licensed before use and that, had he known of unlicensed scraping, he would have required OpenAI to retrain its models. OpenAI declined to comment on the filings.
What’s Next: Ongoing Motions and Potential Outcomes
U.S. District Judge Sidney H. Stein (Southern District of New York) is reviewing motions for summary judgment. The court’s forthcoming rulings will determine whether the companies’ use of news content qualifies as fair use and could set precedent for how AI developers may train models on copyrighted material.
*All factual claims are drawn from the unsealed court documents and public statements referenced in the source material.*
