Full Breakdown
Internal Execs Call AI News Scraping “Astonishing Theft” as Court Docs Unsealed
By Drooid · · How we work
Core Event: Unsealed Admissions in the NYT Copyright Lawsuit
The documents also warned of a “doom loop” in which AI chatbots siphon traffic from news sites, potentially harming the very content supply chain that powers the models.
Background & Context
The lawsuit was filed in December 2023 after *The New York Times* alleged that Microsoft’s Copilot and OpenAI’s ChatGPT were built on millions of the paper’s copyrighted articles without permission or compensation. The case has expanded to include the *Chicago Tribune*, the *Daily News*, the *Center for Investigative Reporting*, the Authors Guild, and numerous individual writers. Both companies have argued that their use of the material qualifies as “fair use,” a defense that judges have so far treated cautiously because the law has not yet settled on AI training practices.
Data & Statistics
- The filing cites more than 3.9 million copies taken from *The Times* and 7.3 million from the *Tribune* and its sister papers for LLM training.
- A separate count shows over 91,692 distinct works from *The Times*, *Daily News* and the *Center for Investigative Reporting* appearing in OpenAI’s mid-training dataset.
- A Common Crawl-derived set contains more than 2 million documents pulled from nytimes.com alone.
- Internal Microsoft data indicate that Copilot reduced click-through rates to *The Times* by up to 93 % compared with traditional Bing search.
- Survey results reveal that 36 % of *Times* subscribers using ChatGPT said they no longer need the newspaper, while 28.6 % of former subscribers are more likely to ask an AI for news.
- Only 1.3 % of the roughly 45,000 Copilot conversations were about current affairs, but ChatGPT, with 1 billion monthly users as of May, registers about 1 million news-related prompts per week.
All figures are drawn from the unredacted court materials filed by the plaintiffs.
Official Statements & Responses
- The Trump administration submitted a brief supporting the tech companies, arguing that restricting LLM development would “thwart creative and scientific progress” and harm U.S. economic competitiveness.
Criticism & Opposition
Publisher counsel contends that the internal admissions prove the companies are “knowingly committing theft,” contradicting the fair-use defense. The attorneys argue that the AI products directly substitute for the original news content, undermining the market for the publishers’ work.
Conflicting Reports & Gaps
The filings present several overlapping but distinct counts of copied material (millions of copies overall, 91,692 specific works, 2 million Common Crawl documents). The lack of a single consolidated total leaves uncertainty about the full scope of the alleged infringement. Additionally, while the plaintiffs provide detailed click-through and survey data, the companies have not publicly responded to those specific metrics.
Verbatim Quotes
- “Throughout this case, defendants insisted that these documents be treated as confidential so that the public could not see them,” — Steven Lieberman, attorney
What’s Next
The presiding judge is expected to rule on whether the case proceeds to trial sometime in 2027. Both sides have filed summary-judgment motions, and further unsealing of evidence may occur as the court evaluates the fair-use arguments.
