AI crawlers and robots.txt:allow AI search, block training
GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and Google-Extended explained, with copy-paste robots.txt rules to keep AI search and opt out of training.

On this page
- What AI crawlers and robots.txt actually are
- Before you start
- Decide what you want each bot to do
- How robots.txt rules are read
- Copy-paste robots.txt examples
- How to edit robots.txt in cPanel
- How to edit robots.txt in WordPress
- Using Cloudflare for AI crawlers
- Optional: block bots at the server
- How to check it worked
- Troubleshooting
- When to ask your host or a developer
- The one message worth forwarding
- Reader questions
- Sources & further reading
The short answer
AI providers publish separate controls for model training, AI search and user-requested fetches. To allow selected AI search bots while opting out of the covered training uses, disallow GPTBot, ClaudeBot and Google-Extended while leaving OAI-SearchBot, Claude-SearchBot, PerplexityBot and Googlebot allowed. Google-Extended also controls covered Gemini grounding. robots.txt is a voluntary instruction, not access control or a universal AI opt-out.
Several major AI providers publish separate controls for training, AI search and user-requested page fetches. You can use robots.txt to ask the covered training crawlers to stay out while allowing selected search bots. These controls are not a universal opt-out from every AI use.
Here is the part that surprises people. Blocking GPTBot does not remove you from ChatGPT search, because OpenAI uses a different bot, OAI-SearchBot, for that, and its docs say "each setting is independent of the others". Blocking GPTBot expresses an opt-out from OpenAI's training use of crawled content. Blocking OAI-SearchBot excludes a site from ChatGPT search answers, although navigational links can still appear.
By the end of this guide you will know which bot does what, have a copy-paste robots.txt for your choice, know how to add it in cPanel, WordPress or Cloudflare, and be able to prove it works.
What AI crawlers and robots.txt actually are
A crawler (or bot) is a program that downloads web pages automatically. Each request carries a user-agent, a short text label that says who is asking, such as GPTBot/1.4.
robots.txt is a plain text file at the root of your site, for example https://example.com/robots.txt, with rules like "this bot may not visit these paths". The standard behind it is RFC 9309, the Robots Exclusion Protocol.
Think of it as a sign on your shop door: "Delivery drivers welcome, survey takers please stay out". Polite visitors follow it, but it is not a lock. RFC 9309 says it directly: "These rules are not a form of access authorization."
Training bots, search bots and user fetchers
Once you see this split, every vendor's docs become easy to read.
- Training crawlers collect pages that may be used to train AI models. Blocking them is the "opt out of training" choice.
- AI search crawlers build an index so an assistant can show and link your pages in answers. Blocking them means fewer chances to be found there.
- User-triggered fetchers visit a page because a person asked the assistant to read it. Some vendors say these may not follow
robots.txt, because a human started the request.
The exact tokens, verified today
At the time of writing (3 October 2026), these are the tokens each vendor documents. A token is the short name you write after User-agent: in the file.
OpenAI:
GPTBot: crawls content "that may be used in training our generative AI foundation models".OAI-SearchBot: "used to surface websites in search results in ChatGPT's search features".ChatGPT-User: visits pages for actions users ask for in ChatGPT; OpenAI says "robots.txt rules may not apply".- For search results, OpenAI says it can take about 24 hours for its systems to adjust after you update
robots.txt.
Anthropic:
ClaudeBot: collects content for model training. If blocked, Anthropic says your future materials "should be excluded from our AI model training datasets".Claude-SearchBot: crawls to improve search results. Blocking it "prevents our system from indexing your content for search optimization".Claude-User: fetches pages when a Claude user asks a question. Blocking it stops Claude retrieving your content for those questions.- Anthropic says it honours
robots.txt, supports the non-standardCrawl-delayrule, and asks you to set rules on every subdomain you want to opt out.
Perplexity:
PerplexityBot: surfaces and links sites in Perplexity search, is "not used to crawl content for AI foundation models", and respectsrobots.txt.Perplexity-User: fetches pages when a user asks, and "generally ignores robots.txt rules".- Perplexity says changes may take up to 24 hours to show.
Google:
Google-Extendedis a "standalone product token" that controls whether content Google crawls may be used to train future Gemini models and for grounding in Gemini Apps and Grounding with Google Search on Vertex AI.- Google says it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search".
- It has no user-agent string of its own. Normal Google crawlers do the fetching, so you will never see it in logs and a server rule that matches it does nothing.
Why this matters: Googlebot crawls for Google Search. Blocking it can damage search visibility, but blocked URLs may still appear in results. Google-Extended is the separate control for the Gemini uses described above.
If your real goal is being quoted in AI answers, read how to get your website cited by ChatGPT, Claude and Google's AI answers. This guide covers only the access rules.
Before you start
The edit is small, but a mistake in robots.txt can hide your whole site from search engines.
- Access to cPanel File Manager, FTP or SSH, or your WordPress admin.
- Your current rules: open
https://example.com/robots.txt(with your domain) and copy the text into a note. - A backup: in File Manager, select
robots.txt, click Copy and save it asrobots.txt.bak. - A decision on what you want (next section).
- About 15 minutes, plus up to a day for vendors to pick up the change.
- Risk: low if you only add AI bot groups; high if you edit the
User-agent: *group or anything forGooglebot.
Decide what you want each bot to do
There is no single right answer; it depends on what matters more to you.
- "AI search traffic, and training is fine." No change is needed if your existing
robots.txtand firewall already allow the relevant bots. Without a matching named group, bots still follow yourUser-agent: *rules. - "Selected AI search traffic, but no training by these providers." Block
GPTBot,ClaudeBotandGoogle-Extended; allowOAI-SearchBot,Claude-SearchBotandPerplexityBot. BlockingGoogle-Extendedalso restricts the covered Gemini grounding uses, not just training. - "Limit known AI access." Block the listed tokens and consider server or Cloudflare rules for detected crawlers. User-triggered fetchers and unidentified bots may still visit; these controls cannot guarantee that a public page stays out of every AI system.
- "Only some public folders should be skipped." Disallow those paths for the relevant bots. Protect genuinely private content with authentication and authorization, not
robots.txt.
A made-up example: Amira runs a small bakery site with her own recipes. She wants people asking ChatGPT for "custom cakes near me" to find her, but she does not want her recipe writing used for training. She picks the second option.
How robots.txt rules are read
A few rules from RFC 9309 explain almost every surprise.
A bot follows only its own group
A group is one or more User-agent: lines followed by Allow: and Disallow: rules. Matching groups for the same token are combined. The User-agent: * group is the fallback when no named group matches.
Heads up: named groups do not inherit your * restrictions. Preserve applicable path restrictions when adding a named group; a group already using Disallow: / blocks every path.
Token matching is case-insensitive, so gptbot works the same as GPTBot.
The longest matching path wins
When an Allow and a Disallow both match a URL, the rule with the longest path wins; on a tie, Allow wins. Paths are case-sensitive, and * means "any characters" while $ means "end of the URL", so Disallow: /*.pdf$ blocks URLs ending in .pdf.
Errors change the meaning
If robots.txt returns a 4xx error such as 404, the RFC says crawlers may access everything. If it returns a 5xx server error, they must assume everything is disallowed. Google pauses crawling for 12 hours in that case, then uses the last good copy for up to 30 days.
The file must be named robots.txt in lowercase, sit at the top level of the host and be UTF-8 plain text; Google ignores content beyond 500 KiB. Each host and protocol needs its own file, so shop.example.com does not inherit rules from example.com.
Copy-paste robots.txt examples
These are minimal example files. Merge the relevant bot groups into your current file; preserve existing restrictions and sitemap lines, and check duplicate groups. Replace the example Sitemap URL with yours, or remove it.
Example 1: Allow AI search, block AI training
This file blocks the training crawlers, welcomes AI search bots and Google Search, and keeps the usual WordPress admin rule in both open groups.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
Disallow: /
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Sitemap: https://example.com/sitemap.xmlQuick tip: on a site without WordPress, replace each pair of wp-admin rules with an empty Disallow: line, preserving the group boundaries.
Example 2: Block the AI tokens covered here
This file asks every AI bot in this guide to stay away and leaves normal search engines alone.
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
Disallow: /
User-agent: *
Disallow:
Sitemap: https://example.com/sitemap.xmlChatGPT-User and Perplexity-User may not obey this, by their vendors' own docs, but listing them still states your preference.
Example 3: Ask selected AI bots to skip one folder
This example asks the five named bots to skip /members/ while allowing other paths. It does not cover every AI service or user-triggered fetcher. Use authentication for private members-only content.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Disallow: /members/
User-agent: *
Disallow:Quick tip: if ClaudeBot is only too busy, Anthropic documents Crawl-delay: 1 under User-agent: ClaudeBot to slow it down instead. Crawl-delay is not in RFC 9309 and Google does not use it.
How to edit robots.txt in cPanel
This works for any site, WordPress or not.
Step 1: Open File Manager and show hidden files
In cPanel open File Manager in the Files section. Open Settings, tick Show Hidden Files (dotfiles) and save, so you can see .htaccess later.
Step 2: Open the document root
The document root is the folder your domain serves files from, usually public_html for the main domain. If unsure, ask your host.
Step 3: Back up or create the file
If robots.txt exists, select it, click Copy and copy it to robots.txt.bak. If not, click + File, type robots.txt in New File Name, check the path and click Create New File.
Step 4: Edit and save
Select robots.txt, click Edit, merge the chosen bot groups with your existing rules, and save. Open https://example.com/robots.txt in a private window to check the complete result.
How to edit robots.txt in WordPress
Your WordPress site may show a robots.txt even though no such file exists on the server.
Virtual file versus physical file
With no real file, WordPress generates a virtual one. Its default is User-agent: *, Disallow: /wp-admin/ and Allow: /wp-admin/admin-ajax.php, and plugins can change it through the robots_txt filter.
The standard WordPress .htaccess only passes a request to WordPress when no real file matches. So once you upload a physical robots.txt, the server sends that file and the virtual one, with any plugin additions, disappears. If a plugin used to add your Sitemap line, add it by hand.
Option A: Upload a physical file
Follow the cPanel steps above. What you see in File Manager is then exactly what bots get.
Option B: Yoast SEO's file editor
In Yoast SEO, go to Yoast SEO > Tools > File editor, click Create robots.txt file if none exists, edit and save. Yoast notes the menu does not appear when file editing is disabled, for example with DISALLOW_FILE_EDIT; use cPanel then.
Heads up: do not use the WordPress option that discourages search engines for this. It affects every search engine, and since WordPress 5.3 it uses a robots meta tag rather than writing Disallow: /.
Using Cloudflare for AI crawlers
If your domain runs through Cloudflare, two tools help, and both are on all plans at the time of writing (3 October 2026).
Managed robots.txt
Cloudflare can generate rules telling known AI crawlers to stay away. If you already have a robots.txt, Cloudflare prepends its rules to yours and serves both as one file; if you have none, it creates one.
To turn it on, open the Security Settings page, filter by Bot traffic, and turn on Set your preference to block training in robots.txt. The result includes a line like Content-signal: search=yes, ai-train=no, use=reference.
Heads up: bots now see more than the file on your server, so always test the live URL.
AI Crawl Control
AI Crawl Control shows which AI services visit and lets you block each crawler at Cloudflare. Open AI Crawl Control, go to the Crawlers tab, and in the Action column select Block. Cloudflare notes that "robots.txt compliance is voluntary", which is why enforcement exists.
Optional: block bots at the server
robots.txt is a request; a server rule is a refusal that answers 403 Forbidden. Use it only for bots you never want.
Some warnings first. A user-agent is text anyone can fake, and the Apache docs say any user-agent technique "can be trivially circumvented". Anthropic warns that blocking that stops its bots reading robots.txt may not give a reliable opt-out, so the rules below still let bots read that file. OpenAI, Anthropic and Perplexity publish their bots' IP ranges as JSON files on their bot pages, so check an odd visitor's IP before blaming the vendor.
Apache or LiteSpeed: .htaccess
In File Manager, select .htaccess in the document root, click Copy to save .htaccess.bak, then click Edit. Paste the block below at the very top, above any # BEGIN WordPress line.
This rule returns 403 to requests whose user-agent contains GPTBot or ClaudeBot, except requests for robots.txt. Serve robots.txt as a physical file for this simple example; a virtual WordPress file can be affected by internal rewrites. Confirm the bot receives 200 for /robots.txt before leaving the block enabled.
RewriteEngine On
RewriteCond %{REQUEST_URI} !^/robots\.txt$
RewriteCond %{HTTP_USER_AGENT} (GPTBot|ClaudeBot) [NC]
RewriteRule .* - [F][NC] ignores upper and lower case and [F] sends 403. Add bots inside the brackets with |, such as (GPTBot|ClaudeBot|PerplexityBot). Apache's docs show SetEnvIfNoCase with Require not env for blocking a bad robot and mod_rewrite as a fallback; the rewrite form is simpler in .htaccess because it exempts robots.txt in one line. If the file already holds an HTTPS redirect, our guide on forcing HTTPS with .htaccess explains how blocks sit together.
Nginx
Nginx does not read .htaccess, so you edit the server config over SSH. Copy the site's config file first, outside the folder Nginx loads configs from.
This map goes in the http block and flags matching user-agents.
map $http_user_agent $block_ai_bot {
default 0;
"~*GPTBot" 1;
"~*ClaudeBot" 1;
}These lines go in your site's server block: they copy the flag, clear it for /robots.txt, and return 403 when it is still set.
set $deny_bot $block_ai_bot;
if ($uri = /robots.txt) {
set $deny_bot 0;
}
if ($deny_bot = 1) {
return 403;
}These commands test the configuration and then reload Nginx; run the second only if the first passes.
sudo nginx -t
sudo nginx -s reloadHow to check it worked
Replace example.com with your domain in every command.
Check 1: The live file
This command prints the file bots actually receive, including anything Cloudflare or a plugin adds.
curl -s https://example.com/robots.txtThis one prints only the status code, which should be 200.
curl -s -o /dev/null -w "%{http_code}\n" https://example.com/robots.txtCheck 2: The server block hits bots only
These commands send a GPTBot user-agent and then a browser-like one, printing each status code.
curl -s -o /dev/null -w "%{http_code}\n" -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot" https://example.com/
curl -s -o /dev/null -w "%{http_code}\n" -A "Mozilla/5.0" https://example.com/Expect 403 then 200. Run the first again with /robots.txt on the end of the URL; it should print 200.
Check 3: Google Search Console
In Search Console, open Settings and then the robots.txt report. It lists the files Google found for your top 20 hosts, when each was checked, the fetch status and any issues. After a fix, select the settings icon next to the file and choose Request a recrawl.
Check 4: Access logs
Ask your host where your raw access log is. These commands count GPTBot lines and show the latest visits from the search and training bots, ignoring case.
grep -i -c "GPTBot" access.log
grep -i -E "OAI-SearchBot|ClaudeBot|Claude-SearchBot|PerplexityBot" access.log | tail -n 20About a day after your change, blocked training bots should mostly request only robots.txt.
Troubleshooting
robots.txt returns 404 or old rules
Symptom: curl prints 404 or the old text. Cause: the file is outside the document root or named Robots.txt. Fix: move it into the document root with a lowercase name; remember a 404 tells bots everything is allowed.
Bots still crawl after the change
Symptom: the live file is right but GPTBot still shows in logs. Cause: caching; RFC 9309 lets crawlers keep the file up to 24 hours, and Google may keep it longer if it cannot refresh. Fix: wait a day, then compare the IP with the vendor's published ranges, since a fake bot ignores your file.
Googlebot blocked by mistake
Symptom: Search Console reports pages "blocked by robots.txt", or traffic falls. Cause: Disallow: / under User-agent: * instead of under the AI group. Fix: restore robots.txt.bak, add only the AI groups, then Request a recrawl.
Lines you never wrote
Symptom: the live file has extra rules. Cause: WordPress serving a plugin-built virtual file, or Cloudflare's managed robots.txt. Fix: upload a physical file and check the Cloudflare Security Settings page. Purge any page cache too; our LiteSpeed Cache guide shows where.
Real visitors get 403 Forbidden
Symptom: you or customers see "403 Forbidden". Cause: the pattern matches more than the bot you meant, or sits in the wrong place. Fix: restore .htaccess.bak, use the exact tokens above and repeat Check 2.
500 Internal Server Error after editing .htaccess
Symptom: the whole site shows a 500 error. Cause: a typo or a directive your host does not allow. Fix: restore .htaccess.bak at once, because a 5xx on robots.txt makes crawlers treat the site as fully disallowed.
When to ask your host or a developer
Ask for help if you run Nginx without SSH access, use a CDN other than Cloudflare, or need bots verified by IP, which is firewall work. For the bigger picture, see AI-powered hosting: what actually changes.
The one message worth forwarding
Copy this to whoever looks after your website:
"Hi, please update our robots.txt so AI search bots can read the site but AI training crawlers cannot. Add one group with User-agent lines for GPTBot, ClaudeBot and Google-Extended and Disallow: /, and keep OAI-SearchBot, Claude-SearchBot, PerplexityBot and Googlebot allowed. Keep our existing rules, back up the old file first, and confirm with curl that our /robots.txt returns 200 with the new rules. No server-level blocks for now. Thanks!"
The takeaway: use each provider’s documented tokens to express your crawling and content-use preferences, preserve existing rules, and verify the live result. Use access controls for private content.
Reader questions
Does blocking GPTBot remove my site from ChatGPT search?
No. OpenAI uses OAI-SearchBot for ChatGPT search and GPTBot for training, and says each setting is independent. Block only GPTBot if you want to opt out of training but stay visible in ChatGPT search.
Will blocking Google-Extended hurt my Google rankings?
Google says Google-Extended does not affect a site's inclusion in Google Search and is not a ranking signal. It controls use of your content for Gemini model training and grounding, not Googlebot crawling for Search.
Does robots.txt keep a page out of Google search results?
No. A disallowed URL can still appear in Google results. To remove a public page, use a noindex meta tag or HTTP header while allowing Googlebot to crawl it and see the rule. Protect private content with authentication.
Can AI bots ignore my robots.txt?
Yes. RFC 9309 says robots.txt is not a form of access authorization, and OpenAI and Perplexity say their user-triggered fetchers may not follow it. Use a server rule or Cloudflare AI Crawl Control if you need enforcement.
How long does a robots.txt change take to work?
OpenAI says about 24 hours for search results and Perplexity says up to 24 hours. Google generally caches robots.txt for up to 24 hours, longer if it cannot refresh the file.
Do subdomains need their own robots.txt?
Yes. Google applies a robots.txt only to the host, protocol and port where it is served, and Anthropic asks you to set rules on every subdomain you want to opt out.
Should I block AI crawlers by IP address instead?
Anthropic warns that IP blocking may not give a reliable opt-out because it stops its bots reading your robots.txt. Start with robots.txt, and use the vendors' published IP ranges mainly to tell real bots from fakes.
Sources & further reading
- OpenAI: Overview of OpenAI crawlers
- Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?
- Perplexity: Perplexity crawlers
- Google Search Central: Google's common crawlers
- Google Search Central: How Google interprets the robots.txt specification
- RFC 9309: Robots Exclusion Protocol
- Cloudflare: Managed robots.txt
- Cloudflare: AI Crawl Control
- Apache HTTP Server: When not to use mod_rewrite
- cPanel Docs: File Manager
Originally published . About our editorial updates.


