AI search visibility

Should you block AI crawlers?

Should you block AI crawlers?

In short. For most businesses, no. A crawler you block cannot cite you, so blocking is a decision to remove yourself from that engine's answers. Blocking makes sense in three specific cases: you license your content, you are a publisher whose business model depends on the click, or the crawler is hammering your server. Everyone else is choosing invisibility to solve a problem they do not have.

The question arrives with more heat than most SEO decisions, because it mixes measurement with principle. It is worth separating the two. Blocking an AI crawler has one concrete, predictable consequence: that system can no longer quote you. Whether that is a loss or a win depends entirely on how you make money. Google's own crawler documentation and OpenAI's bots page both set out what each agent does, and reading them is more useful than any opinion, including this one.

What do these crawlers actually do?

Lumping them together is the root of most bad decisions here. There are three jobs, and they have very different consequences for you.

Training crawlers collect text to train future models. Blocking them affects a model that does not exist yet, and has no effect on answers being generated today.

Retrieval crawlers fetch pages to build the search index an assistant consults when answering. Blocking one of these removes you from current answers, immediately.

User-triggered fetchers load a page because a person asked the assistant about it, right now. Blocking these means a user who explicitly asked about your page gets told the assistant cannot read it.

CrawlerOperatorJobBlocking costs you
GPTBotOpenAITraining and indexPresence in ChatGPT answers
OAI-SearchBotOpenAISearch indexCitations in ChatGPT search
ChatGPT-UserOpenAIUser-triggered fetchA user who asked about your page directly
PerplexityBotPerplexitySearch indexCitations in Perplexity
ClaudeBotAnthropicTraining and retrievalPresence in Claude answers
Google-ExtendedGoogleGemini training useGemini use only, not Search rankings
Applebot-ExtendedAppleApple IntelligenceApple Intelligence answers
CCBotCommon CrawlOpen datasetInclusion in many downstream datasets

Does blocking Google-Extended hurt my rankings?

No, and this is the single most common misunderstanding in the whole topic. Google-Extended is a separate control governing whether your content may be used for Gemini and related generative products. It has no effect on Googlebot's crawling, on indexing, or on how you rank in Google Search.

Blocking Googlebot is what removes you from Search, and nobody does that by accident more than once.

The genuinely ambiguous case is AI Overviews. Those are served inside Google Search and draw on the Search index, so opting out of Google-Extended does not remove you from them. If your goal is not appearing in AI Overviews, the only lever is the nosnippet family of directives, and those also suppress your ordinary snippet, which almost always costs more than it saves.

What does blocking actually cost?

Three things, in rough order of how much they hurt.

You leave the answer, and your competitor does not. Assistants do not report an absence. The user simply reads a recommendation naming three companies, none of which is you, and never learns you exist. This is invisible in your analytics, which is what makes it dangerous.

You lose disproportionately valuable visits. Published analyses consistently find AI-referred visitors converting at higher rates and staying longer than ordinary organic traffic. The volume is small today; the quality is not.

You lose the compounding effect. Being cited in an assistant's answer builds the brand association that makes the next citation more likely. Opting out for a year is not a pause, it is a year of ground given to whoever stayed in.

Set against that, the cost of allowing is close to zero for most sites: some crawl load, and content you already published publicly being read by a machine that credits you for it.

When is blocking the right call?

Three cases, and they are narrower than the discourse suggests.

You license your content. If your archive is an asset you sell, giving it away for free to a system that resells access to it is a straightforward commercial error. Publishers with syndication deals block for exactly this reason, and they are right to.

Your business model is the click itself. An ad-funded publisher earns nothing from a citation without a visit. If every impression must become a session to pay for itself, an assistant that summarises your article and names you is taking your inventory and returning a footnote.

The crawl is a genuine load problem. Some agents crawl aggressively. This is an operational issue with an operational fix: rate-limit or block the specific offending agent, rather than treating it as a policy question.

Notice what is not on that list: disliking AI, worrying vaguely about content theft, or blocking because a competitor did. None of those survives contact with the question 'what do I gain'.

How do you implement it correctly?

robots.txt is the mechanism, and it needs three things understood about it.

It is a request, not enforcement. The major named operators publish that they honour it and independent log analyses generally support that. Nothing technically compels compliance, and scrapers that ignore it will keep ignoring it.

It controls crawling, not indexing. A URL disallowed in robots.txt can still appear in an index if other pages link to it.

It is public. Never use it to hide anything, because the file is a published list of what you did not want found.

The safest default for a business that wants visibility is: allow every named retrieval crawler, allow the user-triggered fetchers, and make a deliberate decision about training-only agents such as CCBot. Our robots.txt generator builds the file with all fifteen agents individually toggleable and explains what each one does, so the decision is explicit rather than inherited from a template you found.

Whatever you decide, verify it. Fetch your own robots.txt after deploying and check your server logs a week later to confirm the agents you allowed are actually arriving.

What about the middle ground?

There is one, and it is where most publishers land once they think it through.

Allow retrieval, block training. Allow OAI-SearchBot, PerplexityBot and the user-triggered fetchers so you appear in live answers, and disallow the pure training collectors such as CCBot. You stay visible in the systems people use today while declining to donate your archive to the next model.

Block selectively by path. Allow crawlers on your public guides and documentation, disallow them on gated research, member content or anything you sell. This is usually a better answer than a site-wide decision, because most sites contain both kinds of content.

What does not work is blocking everything and expecting to remain visible. There is no configuration that gives you both, and the vendors implying otherwise are selling something.

Conclusion

Ask one question before you touch robots.txt: does a citation without a click have value to my business? For a SaaS company, an agency, an e-commerce brand or a consultancy, it plainly does, because being named as the recommendation is most of what marketing is trying to buy. For an ad-funded publisher, it plainly might not. Answer that honestly and the configuration follows in about five minutes. Answer it with a general feeling about AI and you will be surprised in six months by traffic you cannot get back.

Build your robots.txt in two minutes AI visibility tracking tools

Frequently asked questions

Does blocking GPTBot remove me from ChatGPT?

It removes your content from what OpenAI collects for training and from the index behind ChatGPT's browsing, so over time you stop being cited. Blocking OAI-SearchBot specifically has the more immediate effect, because that is the agent that builds the search index ChatGPT consults when answering.

Do AI crawlers respect robots.txt?

The major named operators publish that they do, and independent log analyses generally support it. Nothing enforces it technically, so treat robots.txt as a request that reputable operators honour rather than as an access control.

Can I block AI crawlers on some pages only?

Yes, and it is usually the better answer. Use path-level Disallow rules per user agent to keep your public guides visible while excluding gated research, member content or anything you sell. Most sites contain both kinds of content and do not need a single site-wide decision.

Sources

Every figure in this article traces back to one of these. We link them so you can check the original rather than take our summary of it.

Free tools for this

Everything below runs in your browser, with no signup and nothing uploaded.

Definitions: Citation (AI) · AI Overview · LLM visibility · Zero-click search · Prompt tracking

← All articles · Glossary · Statistics

From reading to choosing

Thirty AI SEO tools compared on verified pricing, capabilities and AI search support.

See the 2026 ranking Find your tool in 60 seconds