How to block AI crawlers on WordPress (and keep Google)
· Witen · WordPress guide
Most advice on blocking AI crawlers ends at robots.txt. That file is a request, not a lock. Well-behaved crawlers read it and comply. Nothing makes the rest comply, and some AI fetchers say plainly that they may not read it at all.
Blocking every AI bot is the wrong answer too. Some of them are how your pages get found and cited when people search with ChatGPT, Perplexity, or Siri. The job is to tell them apart, decide what each one gets, and enforce that decision on your own server.
Three kinds of AI traffic
Training crawlers collect pages to build future AI models. Your writing becomes raw material, and nothing links back to you. GPTBot and CCBot are the clearest examples.
AI search crawlers index pages so an assistant can surface and cite them in answers. They work like a search engine’s crawler. OAI-SearchBot and PerplexityBot are in this group, and OpenAI and Perplexity both say these crawlers are not used for training.
User-triggered fetches happen when a person asks an assistant to read a specific page. ChatGPT-User and Perplexity-User work this way. There is a human on the other end, and the visit usually ends in a summary with your link attached.
A sensible default: block training, keep search
For most sites that publish to be read, the right default is simple. Block the training crawlers. Allow the search crawlers and the user-triggered fetches, so your pages can still appear in AI answers with your name on them.
It is your call. A paywalled publisher may block everything. A documentation site may welcome everything. What matters is making the choice per crawler instead of treating every AI request alike, and Google is not part of the trade: Google-Extended, the robots.txt token for Gemini training, does not affect inclusion or ranking in Google Search.
The AI crawlers Witen recognises
Witen Blocker gives each of these crawlers its own allow or block switch. The suggested policy follows the default above. Purposes come from each operator’s own documentation, listed under Sources.
| Crawler | Owner | Purpose | Suggested policy |
|---|---|---|---|
| GPTBot | OpenAI | Training. Collects content that may be used to train OpenAI’s foundation models. | Block |
| OAI-SearchBot | OpenAI | Search. Surfaces sites in ChatGPT’s search results. Not used for training. | Allow |
| ChatGPT-User | OpenAI | User-triggered. Fetches a page when someone asks ChatGPT to. OpenAI says robots.txt may not apply. | Allow |
| ClaudeBot | Anthropic | Training. Collects content for training. Witen’s ClaudeBot switch also covers Claude-SearchBot and Claude-User. | Allow in Witen, disallow ClaudeBot in robots.txt |
| PerplexityBot | Perplexity | Search. Indexes sites for Perplexity answers. Not used for training. Witen’s switch also covers Perplexity-User. | Allow |
| CCBot | Common Crawl | Training. Builds an open web archive that anyone can download, model builders included. | Block |
| Meta-ExternalAgent | Meta | Training. Crawls for training foundation AI models and improving products. Witen’s switch also covers facebookexternalhit, which draws link previews. | Allow in Witen if you share links on Facebook or Instagram, and disallow meta-externalagent in robots.txt. Otherwise block. |
| Amazonbot | Amazon | Products and training. Improves Amazon products and services, and may train Amazon AI models. | Block |
| Applebot | Apple | Search. Powers search in Spotlight, Siri, and Safari. Training is governed separately by Applebot-Extended. | Allow, and disallow Applebot-Extended in robots.txt |
Three rows deserve a second look. Witen’s ClaudeBot, PerplexityBot, and Meta-ExternalAgent switches each cover a family of agents, so blocking one also blocks its search or preview siblings. Where training and search share a switch, keep the switch on allow and turn training off in robots.txt. Anthropic says its crawlers honour robots.txt, and Meta tells site owners to block its crawlers there.
Some operators publish tokens Witen has no switch for: Applebot-Extended and Google-Extended, which control training use without crawling under their own name; Amazon’s Amzn-SearchBot and Amzn-User; and Meta’s Meta-WebIndexer. Handle those in robots.txt.
How Witen enforces the decision
A user agent is just a string, and anyone can type “Googlebot” into one. So Witen checks two things before applying a crawler policy: the user agent must match a recognised crawler, and the request must come from that operator’s published network ranges. Only then does the switch apply.
A blocked crawler gets a 403 before WordPress builds the page. An allowed crawler passes. An impostor claiming to be Googlebot from someone else’s network does not get Googlebot’s pass. Witen treats it as an ordinary visitor, subject to every other rule you run.
This is not a theoretical concern. Common Crawl warns publicly that other crawlers falsely identify themselves as CCBot. A rule keyed on the user agent alone would hand those impostors whatever you decided for the real one.
Nothing is blocked until you choose. Every recognised crawler starts as allowed, and each switch takes effect as soon as you save.
One limit, stated plainly: a scraper that hides its identity behind a browser user agent is not caught by a crawler-specific rule, because there is no crawler to match. Login lockout, IP rules, and filtering at your server or CDN still apply to it.
Set it up in five steps
- Install and activate Witen Blocker from WordPress.org. Crawler controls work without a Witen account.
- Add your own IP address to Witen’s allowlist before changing anything else. If the site sits behind a proxy or CDN, follow the client-IP guide first. Network verification only works when WordPress sees the real visitor address.
- Open Firewall in Witen’s admin menu, then choose Bots and AI crawlers. The AI group holds most of the table above, but look past it: Applebot sits with search engines, Meta-ExternalAgent with social previews, and CCBot with other crawlers.
- Set each crawler to Allowed or Blocked to match your policy, then save the bot settings.
- Publish the robots.txt rules below, then check Witen’s activity over the next few days. A quiet log can simply mean no matching crawler has visited yet.
Add robots.txt as the polite layer
robots.txt complements Witen. It reaches crawlers Witen has no switch for and turns off training where one switch covers both training and search. This version opts out of training and leaves search and user-triggered fetches open:
# Opt out of AI training. Search crawlers and
# user-triggered fetches stay welcome.
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: meta-externalagent
Disallow: /
User-agent: Amazonbot
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: Google-Extended
Disallow: /Expect a delay. Crawlers cache robots.txt: Meta says changes can take up to 24 hours, and Amazon may use a copy up to 30 days old. A Witen switch applies on the next request.
WordPress serves a generated robots.txt unless a real file exists in the site root, and many SEO plugins let you edit it from the dashboard. One quirk: if your file has rules for Googlebot but none for Applebot, Apple says Applebot follows the Googlebot rules.
Keep a way back in
IP rules can lock you out of your own site. Never test a ban against the address you administer from, and keep the login recovery guide handy. When the crawler policy is set, continue with the WordPress setup guide for two-factor authentication, file scans, and optional shared intelligence.