AI crawler directory

Every major AI platform sends its own crawler, for a different purpose. Here's what each one does, and what blocking or allowing it actually means.

OpenAI

Training

GPTBot

OpenAI

Crawls public content that may be used to train future OpenAI models.

Explore the related checker →

If blocked: Your content is excluded from OpenAI's model training crawls. It has no effect on whether ChatGPT can cite your site in search.

If allowed: Your public content becomes eligible for inclusion in OpenAI's training data crawls.

Search / citation

OAI-SearchBot

OpenAI

Powers live web results and citations shown inside ChatGPT search.

Explore the related checker →

If blocked: Your pages become ineligible to appear as a live citation in ChatGPT search answers.

If allowed: Your pages can be fetched and cited in real time when relevant to a user's question.

On-demand fetch

ChatGPT-User

OpenAI

Fetches a specific page when a user explicitly asks ChatGPT to browse it.

Explore the related checker →

If blocked: ChatGPT cannot open your page even when a user directly pastes your URL and asks about it.

If allowed: ChatGPT can open and read your page on a user's direct request.

Anthropic

Training

ClaudeBot

Anthropic

Anthropic's general-purpose crawler for public web content.

Explore the related checker →

If blocked: Your content is excluded from Anthropic's general crawling.

If allowed: Your public content becomes accessible to Anthropic's crawler.

Search / citation

Claude-SearchBot

Anthropic

Powers live citations when Claude uses web search.

Explore the related checker →

If blocked: Your pages become ineligible for live citation in Claude's web search.

If allowed: Your pages can be fetched and cited when relevant to a user's question.

On-demand fetch

Claude-User

Anthropic

Fetches a page when a user directly asks Claude to read a specific URL.

Explore the related checker →

If blocked: Claude cannot open your page even on a direct user request.

If allowed: Claude can open and read your page on a user's direct request.

Google

Indexing

Googlebot

Google

Google's primary crawler for search indexing — the foundation most AI Overviews and Gemini grounding relies on.

Explore the related checker →

If blocked: Your site drops out of Google Search entirely, and AI features built on search grounding lose access.

If allowed: Your site remains eligible for indexing, ranking and AI-grounded features.

Training

Google-Extended

Google

Controls whether your content can be used to improve Gemini and Vertex AI models, separately from search indexing.

Explore the related checker →

If blocked: Your content is excluded from Gemini/Vertex AI model training while remaining indexable in Search.

If allowed: Your content becomes eligible for use in Gemini and Vertex AI model improvement.

Perplexity AI

Search / citation

PerplexityBot

Perplexity AI

Crawls and indexes content to power live citations inside Perplexity answers.

Explore the related checker →

If blocked: Your pages become ineligible to be cited in Perplexity's answers.

If allowed: Your pages can be indexed and cited when relevant to a user's question.

Apple

Indexing

Applebot

Apple

Powers Siri, Spotlight and Safari's on-device intelligence features by indexing public content.

If blocked: Your site drops out of Siri and Spotlight results and Apple's other search-powered features.

If allowed: Your site remains eligible for indexing by Apple's search and on-device intelligence features.

Training

Applebot-Extended

Apple

Controls whether your content can be used to train Apple's foundation models, separately from Applebot indexing.

If blocked: Your content is excluded from Apple Intelligence model training while remaining eligible for Siri/Spotlight indexing via Applebot.

If allowed: Your content becomes eligible for use in training Apple's foundation models.