Esc
IP / CopyrightCase Closed

Reddit lawsuit alleges Perplexity scraped Google to bypass blocks

Is this a scandal?

No longer — the story has resolved. Noise 41/100, holding steady, across 0 sources.

SCAND-177544as of Methodology
Cite this incident"Reddit lawsuit alleges Perplexity scraped Google to bypass blocks." SCAND.Ai incident SCAND-177544, noise 41/100 as of September 19, 2026. https://scand.ai/scandal/reddit-lawsuit-alleges-perplexity-scraped-google-bypass-blocks
FORECASTForecast, not fact

Courts will likely issue preliminary rulings on indirect scraping liability within six months because the honeypot evidence provides specific factual grounds for discovery unlike previous generalized scraping suits.

41

Noise 41/100 — louder than 99% of tracked AI controversies.

AI-assisted analysis · How we work

Why it matters

The case tests whether AI search engines can legally use intermediary platforms to access restricted content, potentially redefining web scraping liability and data licensing norms.

Key points

  1. Reddit alleges Perplexity and SerpApi scraped Google SERPs to retrieve blocked Reddit content.
  2. A honeypot test post visible only to Google reportedly appeared in Perplexity results within hours.
  3. The complaint claims AI search engines rely on search engine indexes for discovery rather than full web crawling.
  4. SerpApi is named as a co-defendant for allegedly facilitating the indirect data extraction.
  5. Legal experts suggest the case tests liability for accessing restricted content via third-party intermediaries.

The story

Reddit has filed a lawsuit against AI search startup Perplexity and scraping provider SerpApi, alleging they circumvented technical barriers to access copyrighted content. The complaint states that Perplexity retrieved Reddit posts by scraping Google search results rather than crawling Reddit directly, effectively bypassing robots.txt restrictions. Reddit reportedly verified this method using a test post visible only to Google crawlers, which subsequently appeared in Perplexity’s outputs within hours. The legal action claims this practice constitutes unauthorized access and copyright infringement. Industry analysts note the allegations suggest major AI search systems rely heavily on existing search engine indexes for content discovery rather than independent web crawling. Perplexity has previously stated it respects publisher preferences, though the lawsuit challenges the legality of obtaining such content through third-party intermediaries. The case could establish significant precedent regarding indirect data acquisition in the generative AI sector.

Who's involved

Critic
Reddit

Alleges Perplexity circumvented technical barriers and infringed copyright by scraping Google to access blocked content.

Defender
Perplexity

Has previously stated it respects publisher preferences and robots.txt directives regarding content access.

Defender
SerpApi

Named as a co-defendant accused of providing the scraping infrastructure used to allegedly bypass restrictions.

Neutral
Alex Groberman

Analyzes the lawsuit's technical claims to argue AI search visibility depends on traditional SEO authority signals.

Most contested claim

Perplexity intentionally circumvented technical barriers by using Google as a proxy to steal copyrighted content

Biggest open question

Whether reliance on Google SERPs is exclusive or merely supplementary to independent crawling remains unverified by independent audit

Read the full story

How we got here

This dispute represents a recurring pattern in digital intellectual property law where technological workarounds outpace established legal frameworks for data access. Historically, conflicts involving automated data collection have evolved from simple trespass-to-chattels claims to complex debates over the Computer Fraud and Abuse Act (CFAA) and the Digital Millennium Copyright Act (DMCA). Prior precedents often turned on whether publicly accessible data could be legally scraped; however, this case introduces the novel variable of 'indirect access' via authorized intermediaries.

The pattern mirrors earlier litigation involving meta-search engines and news aggregators, where defendants argued they were merely linking to or displaying snippets of content already indexed by others. Courts have previously struggled to distinguish between legitimate indexing and commercial misappropriation when the technical mechanism involves a third party. In the AI era, this pattern has shifted from display rights to training and retrieval augmentation, complicating the fair use analysis. The involvement of infrastructure providers like SerpApi echoes past cases against CDN and hosting providers, testing the boundaries of secondary liability for tools that enable mass data extraction. This lineage suggests a judicial trend toward examining the entire data pipeline rather than isolated acts of copying.

The full story

Reddit has filed a federal lawsuit against AI search startup Perplexity and scraping infrastructure provider SerpApi, alleging that the defendants systematically circumvented technical access controls to harvest copyrighted content. The complaint, filed two days ago, centers on the accusation that Perplexity did not merely ignore Reddit’s robots.txt directives but actively bypassed them by scraping Google Search Engine Results Pages (SERPs) as an intermediary proxy. According to the legal filing, this method allowed Perplexity to access and display Reddit content that had been explicitly blocked from direct AI crawling.

To substantiate these claims, Reddit reportedly conducted a forensic verification test three days prior to the suit's public emergence. Platform engineers created a "honeypot" post configured with permissions that made it visible exclusively to Google’s authorized crawlers while blocking all other user agents. According to SEO analyst Alex Groberman, who analyzed the technical allegations, this restricted test post appeared in Perplexity’s outputs within hours of publication. Groberman argues that because the content was inaccessible via standard web crawling due to the specific permission settings, its appearance in Perplexity confirms the AI system was ingesting data directly from Google’s index or a partner service like SerpApi rather than independently crawling the live web.

The lawsuit names SerpApi as a co-defendant, accusing the company of providing the specific scraping infrastructure necessary to execute this alleged bypass. The legal theory posits that using a third-party API to scrape search results for the purpose of accessing otherwise restricted content constitutes both copyright infringement and a violation of anti-circumvention provisions. This distinguishes the case from standard web scraping disputes by focusing on the supply chain of data acquisition rather than just the end-use model training.

Perplexity has previously maintained publicly that it respects publisher preferences and adheres to robots.txt protocols regarding content access. However, the current legal action challenges this defense by alleging that compliance with robots.txt on the destination site is irrelevant if the AI company accesses the content through a search engine intermediary that has already been granted access. The complaint suggests that respecting technical barriers on the front door does not absolve liability when entering through a backdoor provided by a search aggregator.

Alex Groberman’s analysis provides neutral technical context to the legal allegations, suggesting that the lawsuit inadvertently reveals the operational architecture of modern AI search engines. Groberman notes that the incident demonstrates how AI visibility relies heavily on traditional SEO authority signals, recency, and optimized data from major search indices like Google and Bing, rather than comprehensive independent crawling. This perspective reframes the controversy not just as a legal dispute over theft, but as an exposure of the industry's structural dependency on legacy search infrastructure.

The litigation is proceeding in federal court under US District Judge Paul Engelmayer. According to Reuters Legal, Judge Engelmayer has already rejected most of Perplexity’s bid to dismiss the lawsuit, allowing the core allegations regarding copyright violation and data scraping to move forward. This procedural ruling indicates that the court finds sufficient merit in Reddit’s claims to warrant discovery, despite Perplexity’s arguments for dismissal. The rejection of the motion to dismiss signals that the judiciary is willing to examine whether indirect scraping via SERPs carries the same legal weight as direct unauthorized access.

The sequence of events—from the honeypot test to the filing and subsequent judicial review—establishes a tight timeline of escalation. Reddit’s strategy appears to have been evidence-first, securing technical proof of the alleged circumvention before initiating legal proceedings. By naming SerpApi, Reddit has also expanded the scope of liability to include the toolmakers facilitating the data extraction, potentially creating new precedent for intermediary responsibility in the AI data supply chain. The case now moves toward evidentiary phases where the technical reality of how Perplexity ingests Google SERP data will be scrutinized against its public statements on compliance.

What's confirmed, what's disputed

  • ConfirmedReddit created a test post visible only to Google crawlers which appeared in Perplexity outputs within hours
  • ConfirmedReddit sued Perplexity and SerpApi alleging they scraped Google SERPs to bypass robots.txt restrictions
  • ConfirmedUS District Judge Paul Engelmayer rejected most of Perplexity AI's bid to dismiss the Reddit lawsuit
  • DisputedPerplexity relies heavily on Google's top results and authority domains rather than crawling the full internet independently
  • DisputedSerpApi provided the scraping infrastructure used to allegedly bypass Reddit's restrictions

The strongest case each way

Critic's case

The honeypot test provides empirical proof that Perplexity's access was causally linked to Google's index, demonstrating intentional circumvention of access controls rather than incidental indexing, which validates the copyright infringement claim.

Defender's case

Accessing content via a public search engine index that has legitimately crawled the web does not constitute circumvention, as the AI system is consuming publicly available search results rather than breaching a technical barrier itself.

Times this happened before

  • hiQ v. LinkedIn · 2022Supreme Court declined to hear appeal, leaving Ninth Circuit ruling that scraping public data likely doesn't violate CFAA
  • New York Times v. Microsoft/OpenAI · 2023Ongoing litigation establishing framework for AI training copyright claims

What's at stake

Reddit seeks to establish that AI companies cannot use search engines as laundering mechanisms for restricted content. For Perplexity and SerpApi, the stake is the viability of their current data acquisition model; an adverse ruling could mandate expensive architectural changes or licensing fees. For the broader AI industry, the outcome determines whether 'publicly indexed' equals 'free to scrape' for LLM training. While no specific damages are cited in current sources, the precedent affects every AI firm relying on SERP APIs for real-time grounding. Publishers gain a potential legal shield against indirect scraping, shifting power dynamics in data licensing negotiations.

What we still don't know

  • Whether reliance on Google SERPs is exclusive or merely supplementary to independent crawling remains unverified by independent audit
  • Specific contractual or technical role of SerpApi in the alleged bypass is currently based solely on Reddit's complaint assertions

Join the Discussion

Discuss this story

Community comments coming in a future update

Be the first to share your perspective. Subscribe to comment.

Noise Level

Buzz41?Noise Score (0–100): how loud a controversy is. Composite of reach, engagement, star power, cross-platform spread, polarity, duration, and industry impact — with 7-day decay.
Decay: 79%
Reach
54
Engagement
39
Star Power
30
Duration
100
Cross-Platform
20
Polarity
75
Industry Impact
85

The timeline

  1. 3 days ago

    Reddit conducts honeypot verification test

    Platform created a test post visible only to Google crawlers which reportedly appeared in Perplexity outputs within hours.

  2. SEO analyst highlights lawsuit technical details

    Alex Groberman published analysis linking Reddit's allegations to AI search ranking mechanics and dependency on Google SERPs.

  3. 2 days ago

    Reddit files suit against Perplexity and SerpApi

    Complaint alleges defendants scraped Google search results to bypass Reddit's robots.txt restrictions and access copyrighted content.

The full record

Sources & methodology

Every claim above traces to these primary items. How we score →

Where the sources disagree

In dispute Perplexity intentionally circumvented technical barriers by using Google as a proxy to steal copyrighted content

Established Reddit's honeypot test showed blocked content appearing in Perplexity; Judge Engelmayer denied dismissal of copyright claims based on these allegations

What's being under-reported

Missing perspective from SerpApi and Google. SerpApi's terms of service and technical response to being named as co-defendant are absent, leaving the intermediary liability question one-sided. Google's position on whether its SERPs constitute a 'public' space for AI consumption is also unrepresented, despite being central to the technical allegation. Without these views, the narrative over-indexes on Reddit's forensic framing.

Who changed their mind, and why
  • RedditEscalated from technical blocking to active litigation supported by forensic honeypot evidence (was: Technical restriction via robots.txt and crawler blocking)
  • PerplexityMoved from public assurances of compliance to active legal defense following denial of dismissal motion (was: Publicly stated respect for publisher preferences and robots.txt)

The forecast

Courts will likely issue preliminary rulings on indirect scraping liability within six months because the honeypot evidence provides specific factual grounds for discovery unlike previous generalized scraping suits.

Forecast, not fact — an editorial estimate we score when this resolves.

You're up to date

That's the complete picture as of — nothing more to know right now. We'll update this page the moment it changes.