Full Breakdown
Reddit Files Lawsuit Against Data Scrapers Over Unauthorized Data Use
10/22/2025, 11:07:13 PM
Overview of the Lawsuit
On October 22, 2025, Reddit Inc. filed a lawsuit in the U.S. District Court for the Southern District of New York against four companies: Perplexity AI Inc., SerpApi LLC, Oxylabs UAB, and AWMProxy. The lawsuit alleges that these entities unlawfully scraped Reddit's data through Google search results, circumventing both Reddit's and Google's security measures. Reddit seeks financial damages, a permanent injunction against the defendants, and a ban on the use or sale of any previously scraped data.
Allegations of Data Theft
Reddit's complaint describes the defendants as akin to "would-be bank robbers," asserting that they bypassed technological protections to access and extract valuable user-generated content. The lawsuit claims that during a two-week period in July 2025, the defendants accessed nearly three billion search engine results pages containing Reddit data. The complaint emphasizes that Perplexity, which operates an AI-based search engine, has been a significant beneficiary of this data, allegedly using it to enhance its services without proper authorization.
Background and Context
Historically, the relationship between data scrapers and content providers was viewed as symbiotic. However, as AI companies increasingly relied on large datasets for training their models, this dynamic shifted. Reddit, which boasts over 416 million users weekly, has sought to monetize its data through licensing agreements with companies like Google and OpenAI. Despite these efforts, some companies opted to bypass formal agreements, leading to the current legal action.
Key Figures and Statements
Ben Lee, Reddit's chief legal officer, stated, “AI companies are locked in an arms race for quality human content — and that pressure has fueled an industrial-scale ‘data laundering’ economy.” He highlighted the importance of Reddit's content for AI training, noting that the platform is a prime target due to its vast collection of discussions.
Perplexity's spokesperson, Beejoli Shah, responded to the allegations, asserting that the company had not yet received the lawsuit and would "fight vigorously for users' rights to freely and fairly access public knowledge." Similarly, representatives from SerpApi and Oxylabs have indicated their intention to defend against the claims.
Criticism and Opposition
Critics of the data scraping practices argue that they undermine the rights of content creators and pose significant risks to user privacy. The lawsuit reflects broader concerns within the tech industry regarding the ethical use of data and the need for clearer regulations governing AI training practices.
Conflicting Reports and Gaps
While Reddit's lawsuit claims that Perplexity's citations of Reddit data increased "forty-fold" after a cease-and-desist letter was issued, Perplexity maintains that it has not used Reddit content to train its AI models. This discrepancy highlights the ongoing debate over data ownership and usage rights in the AI sector.
What's Next
The case, assigned Case No. 25-cv-8736, is expected to set important precedents regarding data scraping and AI training rights. As legal battles intensify in the AI industry, the outcome of this lawsuit could significantly impact how companies access and utilize online content.
Verbatim Quotes
- “These Defendants are similar to would-be bank robbers, who, knowing they cannot get into the bank vault, break into the armored truck carrying the cash instead,” — Reddit's lawsuit
- “AI companies are locked in an arms race for quality human content — and that pressure has fueled an industrial-scale ‘data laundering’ economy.” — Ben Lee, Chief Legal Officer, Reddit
- “will always fight vigorously for users’ rights to freely and fairly access public knowledge. Our approach remains principled and responsible as we provide factual answers with accurate AI, and we will not tolerate threats against openness and the public interest.” — Beejoli Shah, Spokesperson, Perplexity AI
This lawsuit underscores the growing tensions between content providers and AI companies, as both sides navigate the complexities of data usage in an increasingly digital landscape.
