Is your website blocking ChatGPT without you knowing?
You can have an excellent website and still never appear in an answer from ChatGPT or Perplexity. The cause often lies in a small text file that almost nobody ever opens: robots.txt. A single line in it is enough to deny an AI service access to your entire site. Below you can read how that works, which crawlers exist and which choice you make deliberately.
What robots.txt does
Robots.txt is a file with instructions for crawlers, the programs that read websites automatically. According to the official standard RFC 9309, the file must be called exactly robots.txt and sit in the root of your domain. You will find it at yourdomain.be/robots.txt.
For each crawler, identified by "User-agent", the file states which parts of the site it may read. A line "Disallow: /" means: nothing. In its introduction to robots.txt, Google stresses two limitations. The file is meant to manage crawler traffic and does not keep a page out of search results. Moreover, reputable crawlers follow the instructions, while other crawlers may not. Robots.txt works as an agreement that crawlers follow voluntarily.
The well-known AI crawlers
The major AI companies usually use several crawlers, each with its own purpose. For an SME, the distinction between training and search is the most important one.
OpenAI describes three names in its overview of crawlers:
- GPTBot collects content for training models. If you block it, you indicate that your content may not be used for training.
- OAI-SearchBot reads sites in order to show them in ChatGPT's search answers. According to OpenAI, sites that block it do not appear in those answers.
- ChatGPT-User fetches a page when a user asks for it. OpenAI writes that robots.txt may not apply here, because the action comes from a user.
Perplexity has two names, according to its documentation on crawlers:
- PerplexityBot searches and links websites in Perplexity's search results and, according to the company, is not used to train models. It respects robots.txt.
- Perplexity-User visits pages when a user asks a question and, according to Perplexity, generally ignores robots.txt, because a user gave the instruction.
Anthropic, the company behind Claude, describes three names in its help article on crawlers:
- ClaudeBot collects content that may contribute to training models.
- Claude-SearchBot reads content to improve Claude's search results. Blocking it can reduce your visibility in those results.
- Claude-User fetches pages when a user asks Claude something.
Google works with a separate name: Google-Extended. According to the list of Google crawlers, it lets you decide whether content Google crawls may be used to train future Gemini models. Google states explicitly that Google-Extended has no effect on your inclusion in Google Search and is not a ranking signal. Whether you appear in AI Overviews depends, according to Google's explanation of AI features, on regular indexing: your page must be in Google Search and be eligible to appear with a snippet.
What blocking means
The consequences differ greatly per crawler:
- If you block a training crawler such as GPTBot, ClaudeBot or Google-Extended, you ask that your content not end up in future models. According to the documentation of those companies, this is separate from your visibility in their search features.
- If you block a search crawler such as OAI-SearchBot, PerplexityBot or Claude-SearchBot, you largely disappear from that service's answers.
- A general rule "User-agent: *" with "Disallow: /" blocks all crawlers that comply, Googlebot included.
That last case occurs on sites that were still in development and where the block stayed in place after launch. Sometimes a plug-in or a security service adds rules automatically. A firewall can also keep AI crawlers out without anything in robots.txt, so checking the file alone is not always enough.
What you choose deliberately
For most SMEs the logical choice is to allow search crawlers, because you want to be found by people looking for a supplier through ChatGPT or Perplexity. Opinions on training crawlers can differ. Anyone who publishes unique content on which their business model depends, such as a publisher or a training provider, may want to keep them out. For an installer, a law firm or a wholesaler, that argument usually carries less weight.
A robots.txt that allows search crawlers and refuses training looks like this:
``` User-agent: OAI-SearchBot Allow: /
User-agent: PerplexityBot Allow: /
User-agent: Claude-SearchBot Allow: /
User-agent: GPTBot Disallow: /
User-agent: ClaudeBot Disallow: /
User-agent: Google-Extended Disallow: / ```
More important than the exact choice is that you make it yourself and write it down, so that a future web developer does not quietly undo it. We review this file in every analysis, together with the firewall settings, for SMEs from Ghent and the rest of Flanders.
What you can do now
- Open yourdomain.be/robots.txt and look for the names in this article and for "Disallow: /".
- Ask your web developer or hosting provider whether a firewall or security service is blocking bots. Also ask whether AI crawlers fall under it.
- Decide per type of crawler, search or training, what you want to allow and record that choice.
Let us check this in the free AI analysis. If your website is being redesigned soon, pass these agreements on to our sister agency hetdesignbureau.be as well.
Sources
- RFC 9309: Robots Exclusion Protocol (IETF)
- Introduction to robots.txt (Google Search Central)
- Overview of OpenAI Crawlers (OpenAI)
- Perplexity Crawlers (Perplexity)
- Does Anthropic crawl data from the web and how can site owners block the crawler? (Claude Help Center)
- Google's common crawlers (Google Search Central)
- AI features and your website (Google Search Central)