Skip to content

AI Search Optimization

AI Crawlers Explained: GPTBot, ClaudeBot, PerplexityBot

Add saadrazaseo.com as a preferred source on Google

AI crawlers do three different jobs: training bots such as GPTBot and ClaudeBot collect content for models, search bots such as OAI-SearchBot and PerplexityBot build AI search results, and user-triggered fetchers load a page when someone asks. Each has its own robots.txt rule, so you can block training while allowing search.

AI bots are now a visible part of robots.txt files: the 2025 Web Almanac SEO chapter found GPTBot named in 4.5 per cent of desktop robots.txt files and ClaudeBot in 3.6 per cent, up from 2.9 per cent and 1.9 per cent a year earlier.

When I check a site’s AI crawler access, the first thing I look at is not the list of bots. It is the reason behind each rule. Many sites copy a blocklist from a forum and accidentally switch off the very bots that could cite them. This guide separates the bots by purpose, shows the exact robots.txt lines, and gives you a decision framework you can apply in ten minutes. It is part of the wider topic of generative engine optimisation.

The three jobs an AI bot can do

Vendors have split their crawlers because one bot doing everything forced a bad choice: allow it all or lose AI search visibility. The split gives you three categories.

  • Training crawlers collect content that may be used to train or improve generative AI models.
  • Search crawlers index pages so an AI product can surface and link to them in answers.
  • User-triggered fetchers visit a page because a person asked the assistant a question or gave it a link. They are not crawling the web in bulk.

The category decides your best rule. Blocking a training bot is a content licensing choice. Blocking a search bot is a visibility choice. Blocking a user-triggered fetcher may not even work, as explained below.

The user agents and what each one does

I checked each of these on the vendor’s own documentation. Vendors rename and add bots often, so confirm against the linked pages before you deploy anything.

Vendor User agent or token Purpose Robots.txt behaviour
OpenAI OAI-SearchBot Powers ChatGPT search features Allow it to appear in ChatGPT search results
OpenAI GPTBot Crawls content for training foundation models Disallow to opt out of training
OpenAI ChatGPT-User User-initiated actions such as browsing and custom GPT actions Rules may not apply, because requests are user driven
Anthropic ClaudeBot Collects web content that contributes to training Honours robots.txt, including Crawl-delay
Anthropic Claude-SearchBot Improves search result quality for users Controlled by its own robots.txt group
Anthropic Claude-User Web access when a person asks Claude directly Controlled by its own robots.txt group
Perplexity PerplexityBot Surfaces and links websites in Perplexity search results Perplexity recommends allowing it
Perplexity Perplexity-User Visits a page when a user’s question needs it Generally ignores robots.txt, as users initiate it
Google Google-Extended (token) Controls use of content for training future Gemini models and for grounding A control token, not a separate crawler
Apple Applebot-Extended (token) Controls use of content to train Apple’s foundation models Does not crawl; disallowing it does not remove you from Apple search

The official sources are OpenAI’s crawler documentation, Anthropic’s crawler help article, Perplexity’s bots guide and Apple’s Applebot support page. OpenAI also lists OAI-AdsBot, which checks landing pages submitted as ChatGPT ads, so it only visits pages you submitted.

Google-Extended and Applebot-Extended are not crawlers

This is the most misunderstood part. Google states that Google-Extended is not a separate crawler. Crawling is done with existing Google user agent strings, and the token is used in a control capacity. In Google’s crawler documentation, it is described as the control for whether content is used to train future Gemini models and for grounding, and it does not affect inclusion in Google Search or act as a ranking signal.

That has a practical consequence for AI Overviews. Google’s own guidance says that AI is built into Search and that the robots.txt control for Googlebot is how site owners manage access. Disallowing Google-Extended will not remove you from AI Overviews or AI Mode, and it will not stop them either. Blocking Googlebot would, but it would also remove you from ordinary results, which is almost never what you want.

Apple works the same way. Apple states that Applebot-Extended does not crawl webpages, that pages disallowing it can still appear in search results, and that its rules are not considered in ranking for Search. Applebot itself powers features such as Spotlight, Siri and Safari, so blocking Applebot has a different effect from blocking the token.

How to allow or block each one in robots.txt

Robots.txt works by groups. Each group starts with a User-agent line and is followed by Disallow or Allow lines. If you are new to the file, start with my guide to how robots.txt works. Here are the patterns I use most.

Block training, allow AI search. This suits publishers and businesses that want citations but not model training.

User-agent: GPTBot

Disallow: /

User-agent: ClaudeBot

Disallow: /

User-agent: Google-Extended

Disallow: /

User-agent: Applebot-Extended

Disallow: /

Search bots such as OAI-SearchBot, Claude-SearchBot and PerplexityBot are left unmentioned, so they follow your general rules. If your general group blocks everything, add explicit groups for them with Allow: /.

Allow everything. Add no AI specific rules and make sure no blanket Disallow: / applies to them.

Block one folder only. Use a path, for example Disallow: /members/, inside the group of the bot concerned. This is the sensible option for paywalled or private sections.

Two cautions. First, a bot follows the most specific group that matches its name, so a bot with its own group ignores the generic User-agent: * group. Second, OpenAI says robots.txt changes take about 24 hours to be picked up, so do not judge the result on the same day.

Why blocking user-triggered fetchers is unreliable

Google’s documentation says user-triggered fetchers generally ignore robots.txt because a user requested the fetch. Perplexity says the same about Perplexity-User, and OpenAI says robots.txt rules may not apply to ChatGPT-User. The logic is that a person asked for that page, so the request is closer to a browser visit than a crawl.

If you truly need to stop these requests, robots.txt will not be enough. You would need access controls such as authentication. For most sites this is the wrong goal, because a person asking an assistant about your page is a prospective visitor.

Verifying that a bot is who it claims to be

Anyone can write GPTBot in a user agent string. Robots.txt is also voluntary, so a bad actor can ignore it. Verification protects your logs and your firewall rules.

  1. Pull the user agent and IP from your server logs.
  2. Compare the IP with the vendor’s published ranges. Anthropic points to a published IP list, and Perplexity publishes JSON files for PerplexityBot and Perplexity-User.
  3. For Google, use reverse and forward DNS or the published IP range files. Google documents both methods.
  4. Treat unverified traffic as an impostor. Rate limit or block it at the firewall without touching the real bot’s rules.

Anthropic also warns that blocking by IP address is unreliable, because it can stop the bot reading your robots.txt in the first place. Use robots.txt for policy and IP data for verification.

Server load and the security layer

AI bots rarely cause trouble on a small business site, but they can on large catalogues or sites on weak hosting. Anthropic supports the Crawl-delay extension for ClaudeBot, for example Crawl-delay: 1. Not every vendor documents support for it, so I would not assume it. Caching, a sensible CDN and rate limits at the server level do more.

The bigger risk I see is the opposite one: a host or firewall security challenge that silently blocks legitimate bots and agents, however perfect your robots.txt looks. When I rebuilt saadrazaseo.com in October 2026, a lead form had been blocked by a host security challenge, and I moved it so it no longer blocks agents. I wrote about that wider issue in making a website work for AI agents. Check your firewall rules whenever you change your robots.txt.

A decision framework

Your situation Training bots Search bots Reasoning
Local service business or small shop that wants leads Allow or block, your choice Allow Citations and visibility matter more than the training question.
Publisher or paid content creator Block Allow You protect licensing value while keeping discoverability.
Site with a private or paid area Block that folder Block that folder Use authentication as well, because robots.txt is not security.
Brand with no strong view Allow Allow Simplest setup and the least risk of accidental invisibility.

My rule is to decide on search visibility first and training second. If you want to be cited, read how to get cited in ChatGPT search and how Perplexity picks sources, because access is only the first step. A related question is whether an llms.txt file adds anything. It does not replace any rule here. And I cannot promise that allowing a bot leads to citations. It only removes one reason you would not get them.

A five-minute audit you can run today

  1. Open yoursite.com/robots.txt and read it fully.
  2. List every AI bot named and note whether it is allowed or blocked.
  3. Check that no blanket rule accidentally blocks search bots you want.
  4. Check that your CDN or firewall does not challenge those bots.
  5. Write down your decision and review it every quarter, because vendors change their bots.

If this feels fiddly on a larger site, it is exactly the kind of work covered in my technical SEO services.

Questions people ask about AI crawlers

Will blocking GPTBot remove me from ChatGPT search?

No. OpenAI treats GPTBot and OAI-SearchBot as independent controls. Disallowing GPTBot opts you out of training, while OAI-SearchBot must be allowed for your site to appear in ChatGPT search results.

Does blocking Google-Extended hurt my Google rankings?

Google says it does not affect inclusion in Google Search and is not a ranking signal. It also does not control AI Overviews, which are managed through Googlebot access.

Does robots.txt delete my content from models that already trained on it?

Robots.txt controls future crawling. I have found no vendor documentation promising that a new rule removes data already collected, so I treat it as a forward looking choice.

How do I know which AI bots visit my site?

Filter your server or CDN logs by user agent name, then verify the IP addresses against the vendor’s published ranges. Analytics tools usually miss bots, so logs are the reliable source.

If you want a second pair of eyes on your robots.txt, firewall rules and AI bot access, ask for a free audit through my contact page.

More AI search guides: start with the generative engine optimization guide, then go deeper: