wikiHow sues OpenAI over alleged unauthorized article scraping
Is this a scandal?
Not yet — an early signal. Noise 46/100, heating up, across 1 source.
Courts will likely scrutinize the post-2023 crawling activity separately from pre-block training, because deliberate circumvention of technical barriers undermines standard fair use defenses regarding transformative purpose.
Noise 46/100 — louder than 99% of tracked AI controversies.
Why it matters
This case tests whether ignoring robots.txt directives constitutes willful infringement, potentially redefining fair use boundaries for AI training data collection.
Key points
- wikiHow alleges OpenAI scraped over 11,000 articles including 1,211 registered copyrights without permission or payment.
- The complaint claims OpenAI continued accessing the site hundreds of thousands of times after robots.txt blocks were implemented in 2023.
- wikiHow asserts ChatGPT outputs serve as market substitutes that directly divert advertising and licensing revenue.
- The lawsuit includes DMCA claims for alleged removal of copyright management information like titles and bylines.
- OpenAI defends its practices as fair use involving publicly available data.
- Plaintiff cites Common Crawl datasets as a specific vector for the alleged unauthorized data ingestion.
The story
wikiHow, Inc. filed a copyright infringement lawsuit against OpenAI in the Southern District of New York on August 21, 2026, alleging unauthorized scraping of over 11,000 instructional articles to train ChatGPT. The complaint claims OpenAI accessed the site hundreds of thousands of times after wikiHow implemented robots.txt blocks for GPTBot and related crawlers starting in 2023. wikiHow asserts that ChatGPT generates competing responses that reproduce article substance, diverting traffic and revenue while violating the Digital Millennium Copyright Act by removing copyright management information. OpenAI stated its models are trained on publicly available data and grounded in fair use. The suit seeks monetary damages and a permanent injunction, joining similar litigation by publishers against AI firms regarding training data practices. This filing specifically highlights alleged circumvention of technical blocking measures as evidence of willful misconduct rather than incidental data collection.
Who's involved
Alleges OpenAI willfully infringed copyrights and violated DMCA by scraping content despite technical blocks and stripping attribution.
Maintains that model training relies on publicly available data and is protected under fair use doctrine.
Reports on the filing and confirms presence of wikiHow content in Common Crawl datasets based on personal expertise.
How the conversation shifted
Polarity (0–100) from the noise pipeline, sampled over time.
Noise Level
The timeline
Case details publicized on social media
Brian Roemmele shares docket information and confirms Common Crawl data presence.
Copyright lawsuit filed in SDNY
wikiHow files complaint alleging infringement of 11,000+ articles and DMCA violations.
wikiHow expands crawler blocking measures
Additional directives added to block OAI-SearchBot and ChatGPT-User agents.
wikiHow implements initial robots.txt blocks
Site begins blocking GPTBot and other OpenAI crawlers via robots.txt directives.
The full record
Sources & methodology
- twitter.com — twitter.com
Every claim above traces to these primary items. How we score →
The forecast
Courts will likely scrutinize the post-2023 crawling activity separately from pre-block training, because deliberate circumvention of technical barriers undermines standard fair use defenses regarding transformative purpose.
Forecast, not fact — an editorial estimate we score when this resolves.
That's the complete picture as of — nothing more to know right now. We'll update this page the moment it changes.
Follow this story
We keep this page current — no need to check back. We'll send the next real change to your inbox, nothing else.
Tracking this story since August 25, 2026.
Join the Discussion
Discuss this story
Community comments coming in a future update
Be the first to share your perspective. Subscribe to comment.