Skip to main content

folkfox

Skip to main content
Skip to content
AI CITATIONS

WikiHow's AI Copyright Lawsuit Argues Something Sharper Than Theft

WikiHow is not arguing that ChatGPT copied its instructions. Its ai copyright lawsuit argues something harder to dismiss: that ChatGPT now answers the exact questions WikiHow built a business answering, at WikiHow's expense.

Quick answerThe WikiHow ai copyright lawsuit against OpenAI, filed 21 August 2026, argues market substitution: ChatGPT allegedly reproduces how-to answers that once sent readers, and revenue, to WikiHow's own site.
Section 01

The ai copyright lawsuit arguing something sharper than copying#

A fox does not need to steal the whole henhouse to empty it. It only needs to keep visiting until the hens stop laying for anyone else. That is the uncomfortable shape of the newest ai copyright lawsuit to land against OpenAI, and it is worth reading closely because its argument is genuinely different from the ones that came before it.

WikiHow filed suit against OpenAI in the U.S. District Court for the Southern District of New York on 21 August 2026, according to Playwire's coverage of the filing, alleging OpenAI scraped more than 11,000 of its how-to articles to train ChatGPT, covering upward of 1,200 registered copyrights. OpenAI's standard defence across this whole wave of suits, including the pending Authors Guild class action, is that training on publicly available text is protected fair use. A Reddit thread from SEO practitioner u/chrismcelroyseo, posted the same week, corroborates the article and copyright counts and adds that the complaint also alleges OpenAI kept crawling after robots.txt explicitly blocked it, and stripped copyright management information from the scraped pages.

Most prior AI copyright suits, including the widely covered Getty Images suit against Stability AI, argued straightforward copying: a model reproduced protected text or imagery closely enough to infringe. WikiHow's complaint leads with a different theory entirely. It argues economic substitution, that ChatGPT now "produce[s] competing how-to content on the same subjects, at a fraction of the time, effort, and cost", directly displacing the traffic and advertising revenue WikiHow's own pages once earned. Copying is about the past. Substitution is about the future the plaintiff says it no longer has.

Section 02

Why substitution changes the whole case#

Courts have spent two years wrestling with whether training a model on copyrighted text is fair use. Substitution sidesteps that fight almost entirely. The question stops being "did the model copy the words" and becomes "did the model's answer replace the reason anyone needed the original page at all".

This is not WikiHow's theory alone. A 2025 research paper, The Economics of AI Training Data: A Research Agenda, tracked training-data licensing deals and found five distinct pricing structures already in use, nearly all of which exclude the original creators of the underlying content from any ongoing compensation once the initial deal is struck. Substitution is the economic mechanism that theory describes in practice: the value moves from the page that answered the question first to the system that answers it last.

Illustrative shape of a publisher's lost search value
An illustrative shape of how AI-answered queries erode a publisher's organic value chain, not a measurement of WikiHow's own figures, which remain undisclosed in the complaint.Organic visits: +100 (running total 100)AI Overviews absorb: -22 (running total 78)Chat answers absorb: -18 (running total 60)Remaining visits: 600255075100Organic visits+100AI Overviews absorb-22Chat answers absorb-18Remaining visits60
An illustrative shape of how AI-answered queries erode a publisher's organic value chain, not a measurement of WikiHow's own figures, which remain undisclosed in the complaint.

This incident is not the first, either. Reddit's own SEO community, tracking the filing the week it landed, described it as OpenAI's twenty-fourth active copyright suit, joining a wave of publisher and author litigation the industry has watched build for two years, including The New York Times' own suit against OpenAI and Microsoft and the consolidated multidistrict copyright litigation, MDL No. 3143 now running in the Southern District of New York, all without a single settled precedent on the substitution question specifically. Every one of these cases is, at heart, a different generative ai copyright lawsuit testing a slightly different theory against the same underlying training practice.

Section 03

What the numbers in the filing actually say#

The case in numbers

Articles allegedly scraped

11k+

More than 11,000 WikiHow how-to articles, per the filing.

Registered copyrights at issue

1200+

Copyright registrations named in the complaint.

Prior OpenAI copyright suits

24

This filing reportedly joins OpenAI's 24th active copyright case.

The scale behind the filing
The scale behind the filingBar chart comparing 11000 articles allegedly scraped against 1200 registered copyrights namedArticles allegedly scraped: 11000Copyrights registered: 120015000100005000011000Articles allegedly scraped1200Copyrights registered
More than 11,000 articles allegedly scraped, covering upward of 1,200 registered copyrights, a ratio suggesting most articles are protected individually rather than under a handful of blanket registrations.

WikiHow's complaint reportedly alleges OpenAI's crawlers kept visiting after robots.txt disallowed them, which if proven would matter well beyond this one case. A US appeals court recently held, in a separate dispute covered in folkfox's own reporting on agentic browsing, that a user-directed agent legally counts as the user acting, not the platform. That ruling was about consent to access on a user's behalf. It says nothing about a crawler that keeps returning after a site has explicitly declared it unwelcome, which is closer to what WikiHow alleges here.

Section 04

The ai crawler access control question every publisher now faces#

ai copyright lawsuit: the vixen guards a garden gate deciding which visitor may pass
A gate is only a boundary if something actually stops at it.

robots.txt was never built to be an enforceable legal barrier. RFC 9309, the formal specification the file follows, describes it as advisory: a crawler is meant to fetch and honour it, but nothing in the standard stops one that simply chooses not to. Google's own developer documentation is equally direct that robots.txt is a request, not an access-control mechanism, which is precisely the gap WikiHow's complaint is trying to close with a lawsuit instead of a text file.

That gap is why the fox in this story does not just leave a scent at the gate and hope. WikiHow's complaint treats a crawler that ignores robots.txt as trespass rather than an oversight, which reframes ai crawler access control from a technical setting into a legal exhibit. Every publisher weighing whether to block, allow, or licence AI crawlers is now watching this case for the answer to a question their engineering team cannot settle alone: does robots.txt actually mean anything once it is written down, or does it only work on crawlers polite enough to prowl no further once asked.

ChatGPT now produces competing how-to content on the same subjects, at a fraction of the time, effort, and cost.
WikiHow's complaint, as reported by Playwire
Section 05

What content teams should do this quarter#

Waiting for a court to settle the substitution question is not a content strategy, it is a bet on someone else's litigation timeline. A more useful response starts by measuring, not guessing, how much of a site's own traffic is already quietly slipping into a warren no analytics dashboard shows by default.

A four-step audit any content team can run this month
Check crawler logs

Confirm which AI crawlers are actually visiting, and whether any keep returning after a robots.txt disallow, exactly the pattern WikiHow alleges.

Measure query overlap

Identify which of your highest-traffic informational queries are already showing AI Overviews or chat answers ahead of your listing.

Decide, deliberately, per content type

Choose block, allow or licence per content category rather than site-wide, since a how-to guide and a product page carry very different substitution risk.

Track citations, not just clicks

Start measuring how often your content is cited inside an AI answer, since that is the only visibility metric substitution leaves behind.

does ai violate copyrights is the wrong first question for most content teams to spend their energy on, because it will be litigated for years regardless of what any one publisher decides this quarter. The more useful question is whether a given piece of content still earns a visit once an AI system can answer the same query directly, and that is measurable today, in your own analytics, without waiting for a verdict. A April 2026 citation-absorption study, tracked in this newsroom's own weekly evidence review, found the same substitution pattern WikiHow alleges showing up in raw citation data months before any lawsuit made the argument in court.

This is precisely the audit folkfox runs for SEO and GEO clients already: crawler-log review, query-overlap measurement, and a citation-tracking layer that treats AI visibility as a metric worth reporting on its own, not a footnote under organic traffic. Content strategy now has to account for a reader who may never click through at all, and that shift did not start with this lawsuit. It only became impossible to ignore because of it. Whatever a court eventually decides about this specific ai copyright lawsuit, the underlying traffic pattern it describes is already measurable, today, in every content team's own logs.

Questions

Frequently asked questions#

Does AI violate copyrights?

It depends on the claim. US courts are still deciding whether training a model on copyrighted text is fair use. WikiHow's case argues something narrower and newer: that an AI system's answers can violate a publisher's interests through market substitution, even without directly copying the original wording.

What is the WikiHow AI copyright lawsuit about?

WikiHow's ai copyright lawsuit alleges OpenAI scraped over 11,000 of its how-to articles to train ChatGPT, then produced competing answers to the same questions, displacing the traffic and revenue WikiHow's pages once earned, alongside claims of continued crawling after a robots.txt block.

What is market substitution in an AI copyright case?

Market substitution is the argument that an AI system's output replaces the reason a reader needed the original copyrighted work, harming the creator commercially even if the AI's wording differs from the source text.

How common are AI copyright infringement cases now?

WikiHow's suit is reported as OpenAI's twenty-fourth active copyright case, part of a broader wave of litigation from publishers, authors and rights holders that has built steadily since 2023 without a single settled precedent on market substitution specifically.

Can a website legally stop AI crawlers?

A site can disallow AI crawlers in robots.txt, but that file is a request, not a technical barrier, and crawler compliance is voluntary. Ai crawler access control in practice usually means combining robots.txt with server-level blocking for crawlers that ignore it.

Should publishers block or allow AI crawlers?

There is no single right answer. The decision depends on whether a licensing deal exists, whether the content type is high-value informational content most at risk of substitution, and whether the publisher values AI-answer citations enough to trade some direct traffic for visibility.

Keep reading

Read more on this topic#

Ready to measure your AI visibility, not just your rankings?

folkfox builds GEO and content strategy that accounts for readers who never click through, tracking citations as a real metric, not an afterthought.