WikiHow's AI Copyright Lawsuit Argues Something Sharper Than Theft
WikiHow is not arguing that ChatGPT copied its instructions. Its ai copyright lawsuit argues something harder to dismiss: that ChatGPT now answers the exact questions WikiHow built a business answering, at WikiHow's expense.
By Katie Delaney · 2026-08-28 · 9 min read
The ai copyright lawsuit arguing something sharper than copying#
A fox does not need to steal the whole henhouse to empty it. It only needs to keep visiting until the hens stop laying for anyone else. That is the uncomfortable shape of the newest ai copyright lawsuit to land against OpenAI, and it is worth reading closely because its argument is genuinely different from the ones that came before it.
WikiHow filed suit against OpenAI in the U.S. District Court for the Southern District of New York on 21 August 2026, according to Playwire's coverage of the filing, alleging OpenAI scraped more than 11,000 of its how-to articles to train ChatGPT, covering upward of 1,200 registered copyrights. OpenAI's standard defence across this whole wave of suits, including the pending Authors Guild class action, is that training on publicly available text is protected fair use. A Reddit thread from SEO practitioner u/chrismcelroyseo, posted the same week, corroborates the article and copyright counts and adds that the complaint also alleges OpenAI kept crawling after robots.txt explicitly blocked it, and stripped copyright management information from the scraped pages.
Most prior AI copyright suits, including the widely covered Getty Images suit against Stability AI, argued straightforward copying: a model reproduced protected text or imagery closely enough to infringe. WikiHow's complaint leads with a different theory entirely. It argues economic substitution, that ChatGPT now "produce[s] competing how-to content on the same subjects, at a fraction of the time, effort, and cost", directly displacing the traffic and advertising revenue WikiHow's own pages once earned. Copying is about the past. Substitution is about the future the plaintiff says it no longer has.
Why substitution changes the whole case#
Courts have spent two years wrestling with whether training a model on copyrighted text is fair use. Substitution sidesteps that fight almost entirely. The question stops being "did the model copy the words" and becomes "did the model's answer replace the reason anyone needed the original page at all".
Copying: did the model reproduce protected text?
Courts examine output for close textual similarity to the original copyrighted work.
Substitution: did the answer replace the visit?
Courts examine whether the model's answer satisfies the exact need the original page existed to serve, regardless of wording.
This is not WikiHow's theory alone. A 2025 research paper, The Economics of AI Training Data: A Research Agenda, tracked training-data licensing deals and found five distinct pricing structures already in use, nearly all of which exclude the original creators of the underlying content from any ongoing compensation once the initial deal is struck. Substitution is the economic mechanism that theory describes in practice: the value moves from the page that answered the question first to the system that answers it last.
This incident is not the first, either. Reddit's own SEO community, tracking the filing the week it landed, described it as OpenAI's twenty-fourth active copyright suit, joining a wave of publisher and author litigation the industry has watched build for two years, including The New York Times' own suit against OpenAI and Microsoft and the consolidated multidistrict copyright litigation, MDL No. 3143 now running in the Southern District of New York, all without a single settled precedent on the substitution question specifically. Every one of these cases is, at heart, a different generative ai copyright lawsuit testing a slightly different theory against the same underlying training practice.
What the numbers in the filing actually say#
Articles allegedly scraped
More than 11,000 WikiHow how-to articles, per the filing.
Registered copyrights at issue
Copyright registrations named in the complaint.
Prior OpenAI copyright suits
This filing reportedly joins OpenAI's 24th active copyright case.
WikiHow's complaint reportedly alleges OpenAI's crawlers kept visiting after robots.txt disallowed them, which if proven would matter well beyond this one case. A US appeals court recently held, in a separate dispute covered in folkfox's own reporting on agentic browsing, that a user-directed agent legally counts as the user acting, not the platform. That ruling was about consent to access on a user's behalf. It says nothing about a crawler that keeps returning after a site has explicitly declared it unwelcome, which is closer to what WikiHow alleges here.
The ai crawler access control question every publisher now faces#

robots.txt was never built to be an enforceable legal barrier. RFC 9309, the formal specification the file follows, describes it as advisory: a crawler is meant to fetch and honour it, but nothing in the standard stops one that simply chooses not to. Google's own developer documentation is equally direct that robots.txt is a request, not an access-control mechanism, which is precisely the gap WikiHow's complaint is trying to close with a lawsuit instead of a text file.
That gap is why the fox in this story does not just leave a scent at the gate and hope. WikiHow's complaint treats a crawler that ignores robots.txt as trespass rather than an oversight, which reframes ai crawler access control from a technical setting into a legal exhibit. Every publisher weighing whether to block, allow, or licence AI crawlers is now watching this case for the answer to a question their engineering team cannot settle alone: does robots.txt actually mean anything once it is written down, or does it only work on crawlers polite enough to prowl no further once asked.
ChatGPT now produces competing how-to content on the same subjects, at a fraction of the time, effort, and cost.
What content teams should do this quarter#
Waiting for a court to settle the substitution question is not a content strategy, it is a bet on someone else's litigation timeline. A more useful response starts by measuring, not guessing, how much of a site's own traffic is already quietly slipping into a warren no analytics dashboard shows by default.
Confirm which AI crawlers are actually visiting, and whether any keep returning after a robots.txt disallow, exactly the pattern WikiHow alleges.
Identify which of your highest-traffic informational queries are already showing AI Overviews or chat answers ahead of your listing.
Choose block, allow or licence per content category rather than site-wide, since a how-to guide and a product page carry very different substitution risk.
Start measuring how often your content is cited inside an AI answer, since that is the only visibility metric substitution leaves behind.
does ai violate copyrights is the wrong first question for most content teams to spend their energy on, because it will be litigated for years regardless of what any one publisher decides this quarter. The more useful question is whether a given piece of content still earns a visit once an AI system can answer the same query directly, and that is measurable today, in your own analytics, without waiting for a verdict. A April 2026 citation-absorption study, tracked in this newsroom's own weekly evidence review, found the same substitution pattern WikiHow alleges showing up in raw citation data months before any lawsuit made the argument in court.
This is precisely the audit folkfox runs for SEO and GEO clients already: crawler-log review, query-overlap measurement, and a citation-tracking layer that treats AI visibility as a metric worth reporting on its own, not a footnote under organic traffic. Content strategy now has to account for a reader who may never click through at all, and that shift did not start with this lawsuit. It only became impossible to ignore because of it. Whatever a court eventually decides about this specific ai copyright lawsuit, the underlying traffic pattern it describes is already measurable, today, in every content team's own logs.
Frequently asked questions#
Does AI violate copyrights?
It depends on the claim. US courts are still deciding whether training a model on copyrighted text is fair use. WikiHow's case argues something narrower and newer: that an AI system's answers can violate a publisher's interests through market substitution, even without directly copying the original wording.
What is the WikiHow AI copyright lawsuit about?
WikiHow's ai copyright lawsuit alleges OpenAI scraped over 11,000 of its how-to articles to train ChatGPT, then produced competing answers to the same questions, displacing the traffic and revenue WikiHow's pages once earned, alongside claims of continued crawling after a robots.txt block.
What is market substitution in an AI copyright case?
Market substitution is the argument that an AI system's output replaces the reason a reader needed the original copyrighted work, harming the creator commercially even if the AI's wording differs from the source text.
How common are AI copyright infringement cases now?
WikiHow's suit is reported as OpenAI's twenty-fourth active copyright case, part of a broader wave of litigation from publishers, authors and rights holders that has built steadily since 2023 without a single settled precedent on market substitution specifically.
Can a website legally stop AI crawlers?
A site can disallow AI crawlers in robots.txt, but that file is a request, not a technical barrier, and crawler compliance is voluntary. Ai crawler access control in practice usually means combining robots.txt with server-level blocking for crawlers that ignore it.
Should publishers block or allow AI crawlers?
There is no single right answer. The decision depends on whether a licensing deal exists, whether the content type is high-value informational content most at risk of substitution, and whether the publisher values AI-answer citations enough to trade some direct traffic for visibility.
Read more on this topic#
Agentic browsing just won its first appeal
A related but distinct legal question: what an AI agent is allowed to do once a user authorises it.
Read the pieceChatGPT Citations and the Week Reddit Nearly Vanished
Another robots.txt story, this time about what happens when a platform's own crawler rules shift overnight.
Read the pieceYour robots.txt is a request, and the fetchers know it
The exact enforcement gap in robots.txt that WikiHow's complaint now puts in front of a federal judge.
Read the pieceOpenAI's Own Agents Formed a Swarm. Here Is What Agentic AI Security Missed.
A different OpenAI story from the same week, on how far an unmonitored system can go unnoticed.
Read the pieceReady to measure your AI visibility, not just your rankings?
folkfox builds GEO and content strategy that accounts for readers who never click through, tracking citations as a real metric, not an afterthought.