Esc

wikiHow sues OpenAI over alleged unauthorized article scraping

Is this a scandal?

Not yet — an early signal. Noise 46/100, heating up, across 1 source.

SCAND-212877as of Methodology
Cite this incident"wikiHow sues OpenAI over alleged unauthorized article scraping." SCAND.Ai incident SCAND-212877, noise 46/100 as of August 25, 2026. https://scand.ai/scandal/wikihow-sues-openai-alleged-unauthorized-scraping
FORECASTForecast, not fact

Courts will likely scrutinize the post-2023 crawling activity separately from pre-block training, because deliberate circumvention of technical barriers undermines standard fair use defenses regarding transformative purpose.

46

Noise 46/100 — louder than 99% of tracked AI controversies.

AI-assisted analysis · How we work

Why it matters

This case tests whether ignoring robots.txt directives constitutes willful infringement, potentially redefining fair use boundaries for AI training data collection.

Key points

  1. wikiHow alleges OpenAI scraped over 11,000 articles including 1,211 registered copyrights without permission or payment.
  2. The complaint claims OpenAI continued accessing the site hundreds of thousands of times after robots.txt blocks were implemented in 2023.
  3. wikiHow asserts ChatGPT outputs serve as market substitutes that directly divert advertising and licensing revenue.
  4. The lawsuit includes DMCA claims for alleged removal of copyright management information like titles and bylines.
  5. OpenAI defends its practices as fair use involving publicly available data.
  6. Plaintiff cites Common Crawl datasets as a specific vector for the alleged unauthorized data ingestion.

The story

wikiHow, Inc. filed a copyright infringement lawsuit against OpenAI in the Southern District of New York on August 21, 2026, alleging unauthorized scraping of over 11,000 instructional articles to train ChatGPT. The complaint claims OpenAI accessed the site hundreds of thousands of times after wikiHow implemented robots.txt blocks for GPTBot and related crawlers starting in 2023. wikiHow asserts that ChatGPT generates competing responses that reproduce article substance, diverting traffic and revenue while violating the Digital Millennium Copyright Act by removing copyright management information. OpenAI stated its models are trained on publicly available data and grounded in fair use. The suit seeks monetary damages and a permanent injunction, joining similar litigation by publishers against AI firms regarding training data practices. This filing specifically highlights alleged circumvention of technical blocking measures as evidence of willful misconduct rather than incidental data collection.

Who's involved

Critic
wikiHow, Inc.

Alleges OpenAI willfully infringed copyrights and violated DMCA by scraping content despite technical blocks and stripping attribution.

Defender
OpenAI

Maintains that model training relies on publicly available data and is protected under fair use doctrine.

Neutral
Brian Roemmele

Reports on the filing and confirms presence of wikiHow content in Common Crawl datasets based on personal expertise.

How the conversation shifted

the split has narrowed

Polarity (0–100) from the noise pipeline, sampled over time.

Join the Discussion

Discuss this story

Community comments coming in a future update

Be the first to share your perspective. Subscribe to comment.

Noise Level

Buzz46?Noise Score (0–100): how loud a controversy is. Composite of reach, engagement, star power, cross-platform spread, polarity, duration, and industry impact — with 7-day decay.
Decay: 100%
Reach
49
Engagement
71
Star Power
45
Duration
20
Cross-Platform
20
Polarity
50
Industry Impact
50

The timeline

  1. Case details publicized on social media

    Brian Roemmele shares docket information and confirms Common Crawl data presence.

  2. Copyright lawsuit filed in SDNY

    wikiHow files complaint alleging infringement of 11,000+ articles and DMCA violations.

  3. wikiHow expands crawler blocking measures

    Additional directives added to block OAI-SearchBot and ChatGPT-User agents.

  4. wikiHow implements initial robots.txt blocks

    Site begins blocking GPTBot and other OpenAI crawlers via robots.txt directives.

The full record

Sources & methodology

Every claim above traces to these primary items. How we score →

The forecast

Courts will likely scrutinize the post-2023 crawling activity separately from pre-block training, because deliberate circumvention of technical barriers undermines standard fair use defenses regarding transformative purpose.

Forecast, not fact — an editorial estimate we score when this resolves.

You're up to date

That's the complete picture as of — nothing more to know right now. We'll update this page the moment it changes.

Follow this story

We keep this page current — no need to check back. We'll send the next real change to your inbox, nothing else.

Tracking this story since August 25, 2026.