Drooid Logo
Back to story perspectives

Full Breakdown

Reddit Files Lawsuit Against Perplexity AI and Data Scrapers for Content Theft

10/25/2025, 2:07:30 AM

Allegations of Industrial-Scale Data Scraping

On October 22, 2025, Reddit filed a federal lawsuit in the U.S. District Court for the Southern District of New York against Perplexity AI and three data-scraping firms: SerpApi, Oxylabs UAB, and AWMProxy. The lawsuit accuses these entities of orchestrating an “industrial-scale” scheme to unlawfully scrape Reddit's user-generated content for commercial gain. Reddit claims that the defendants bypassed its security measures by harvesting data through Google search results, thus undermining existing licensing agreements with companies like Google and OpenAI.

The core of Reddit's complaint likens the defendants to “would-be bank robbers,” alleging that they turned to intermediaries to access its data after being blocked from direct scraping. The lawsuit details that in a two-week period in July 2025, the defendants accessed nearly three billion pages of Reddit content through Google, which they then allegedly sold to AI companies.

Background of Legal Actions

This lawsuit is not Reddit's first attempt to protect its data. In June 2025, Reddit filed a similar suit against Anthropic, another AI firm, for unauthorized use of its data. Reddit's legal strategy appears to be aimed at establishing a protective framework around its valuable user-generated content, which it views as a critical asset. The company has previously signed lucrative licensing agreements with major AI players, recognizing the financial potential of its data.

Perplexity's Defense

Perplexity, founded by Aravind Srinivas, has denied all allegations of data theft, asserting that it does not train large language models and operates as an “application-layer” company. The firm claims it lawfully accesses publicly available Reddit data and that its use is limited to summarizing discussions and providing citations for transparency. In a public statement, Perplexity characterized Reddit's lawsuit as a “show of force” in ongoing negotiations with major AI developers, arguing that Reddit is attempting to extort payments through aggressive legal measures.

Implications for the AI Industry

The outcome of this lawsuit could have significant implications for the AI industry and the broader legal landscape regarding data ownership. If Reddit prevails, it could solidify its content as a protected commercial asset, potentially reshaping how AI companies acquire data. Conversely, a victory for Perplexity could legitimize current scraping practices, undermining the emerging data-licensing economy that platforms like Reddit are trying to establish.

Criticism and Opposition

Critics of Reddit's approach argue that its content, once indexed by Google, should be considered public data. Perplexity and the other defendants maintain that they are not engaging in theft but rather utilizing publicly accessible information. This legal battle raises fundamental questions about the nature of data ownership in the digital age, particularly as AI technologies increasingly rely on user-generated content.

What's Next

As the lawsuit progresses, it will likely test the boundaries of the Digital Millennium Copyright Act (DMCA) and the legal definitions of public versus proprietary data. The case, identified as Reddit Inc. v. SerpApi LLC (25-cv-08736), is poised to influence future interactions between social media platforms and AI companies, potentially setting a precedent for how data is accessed and monetized in the evolving digital landscape.

Verbatim Quotes

  • “In a very real sense, these Defendants are similar to would-be bank robbers, who, knowing they cannot get into the bank vault, break into the armored truck carrying the cash instead.” — Reddit Complaint
  • “ “In any case, we won’t be extorted, and we won’t help Reddit extort Google, even if they’re our (huge) competitor.” — Perplexity Statement
  • “AI companies are locked in an arms race for quality human content — and that pressure has fueled an industrial-scale ‘data-laundering’ economy,” — Ben Lee, Reddit's Chief Legal Officer
  • “no company should claim ownership of public data that does not belong to them,” — Denas Grybauskas, Oxylabs Chief Governance and Strategy Officer