Site index for AI agents: /llms.txt. Documentation pages under /docs are also served as markdown at the same URL plus .md (e.g. /docs/geo-audit.md) or via an Accept: text/markdown request.

Which AI Crawlers Visit Your Website, and What Are They Doing?

A log line labelled GPTBot or PerplexityBot is not a citation, and until verified it is not proof the request came from that provider. Classify each request by role, verify its origin against the provider's published method, confirm the page was served, and keep crawler evidence separate from answer evidence.

Harsh Songra, updated

AI companies reach a website for four different jobs, and a single “AI bot traffic” number cannot tell you which one you are looking at.

  • Search discovery. A crawler finds and indexes pages for a search product.
  • User-triggered retrieval. An agent fetches a page while answering a specific person.
  • Training or model improvement.A crawler gathers material under the provider's training controls.
  • Ordinary search and utility crawling. Conventional indexes and platform services fetch the page and may indirectly support AI experiences.

A verified visit proves the provider could access that URL at that time. It does not prove the URL was indexed, retrieved for a later prompt, or cited in an answer.

Four events teams often mix up

EventWhat it meansStrong evidenceWhat it does not prove
AccessA crawler successfully fetched the URL.Verified request, usable status, correct content.Indexing, retrieval or citation.
Discovery / indexingA system knows the page exists and may store it.Provider console, index evidence or repeated verified discovery.That the page will be used for a specific answer.
RetrievalThe page was selected or fetched for a query or agent action.Provider trace, verified user-triggered fetch or answer-time evidence.That the answer cited or relied on it.
Answer selectionThe answer mentioned or cited the page or brand.Saved answer, citation URL, prompt, engine, region and timestamp.That a prior crawl alone caused the selection.

Crawler and control reference

Provider names and policies change without notice. Verify against the linked documentation. Last checked .

ProviderAgent or tokenOperational roleWhat it doesMeasurement note
OpenAIOAI-SearchBotSearch discoverySurfaces websites in ChatGPT's search results, summaries and links.Verify against OpenAI's published ranges.
OpenAIChatGPT-UserUser-triggered retrievalFetches a page because a person or a custom GPT asked for it.OpenAI states robots.txt rules may not apply. A fetch is not a visible citation.
OpenAIGPTBotTraining / improvementCrawls content that may be used to train foundation models.Keep entirely separate from search visibility.
PerplexityPerplexityBotSearch discoveryBuilds and refreshes Perplexity's search index.Perplexity states it is not used for foundation-model training. Verify against published ranges.
PerplexityPerplexity-UserUser-triggered retrievalFetches pages in response to user actions.Perplexity states these requests generally ignore robots.txt. Access is not citation.
AnthropicClaude-SearchBotSearch discoverySupports search quality and result discovery for Claude.The role is distinct from training collection.
AnthropicClaude-UserUser-triggered retrievalFetches pages when a user asks Claude to access them.Check the status and the exact path served.
AnthropicClaudeBotTraining / improvementCollects material for model development under Anthropic's controls.Apply a separate robots policy if the decision differs from search.
GoogleGooglebotSearch discoveryBuilds Google Search's index, which can supply search and AI experiences.Use Google's reverse-and-forward DNS method.
GoogleGoogle-ExtendedControl tokenControls specified Gemini training and grounding uses.Not a separate HTTP user agent. It should not appear as a requester.
AppleApplebotSearch / platform discoverySupports Apple search products and suggestions.Verify using Apple's documented DNS method.
AppleApplebot-ExtendedControl tokenControls specified generative-AI uses of Applebot-crawled content.It does not crawl as a separate bot.
MicrosoftBingbotSearch discoveryBuilds Bing's index, which can support Microsoft search experiences.Verify using Bing's documented method.

Primary documentation: OpenAI crawlers, Perplexity bots, Anthropic web crawlers, Google common crawlers, About Applebot.

Why a user-agent name is not proof

A user-agent header is text the requester supplies. Anyone can copy a crawler name into it. AirPulse has observed requests labelled PerplexityBot probing environment and configuration file paths, which is scanner behaviour, not search discovery.

Label unverified records user-agent-labelled. Upgrade to verifiedonly when the source passes the provider's documented method. Where no reliable method exists, classify as not verifiable rather than guessing the vendor.

How to verify a crawler visit

  1. Keep the raw timestamp, timezone, source IP, full user-agent, host, path, status and response content type.
  2. Identify the provider's official crawler documentation and the role assigned to that agent.
  3. Verify the source against published IP ranges or the provider's reverse-and-forward DNS method. OpenAI publishes ranges for OAI-SearchBot, ChatGPT-User and GPTBot; Perplexity publishes them for PerplexityBot and Perplexity-User; Google and Apple document DNS verification instead.
  4. Confirm the exact URL returned a usable response. A redirect loop, block page, 403, empty shell or 5xx is not successful access.
  5. Classify the event as verified, failed verification, or not verifiable.
  6. Correlate verified access with sitemap discovery, index evidence, prompt runs, citations and outcomes without claiming causality.

What happened on AirPulse's first cited research page

AirPulse examined the page How do I prove that a GEO or AEO change improved AI visibility? and the exact non-branded query with the same wording, over a frozen window. This is an observed sequence, not proof that a crawl caused the citation.

DateObserved event
Jul 27, 2026The page went live. The Research hub link, XML sitemap entry and llms.txt entry shipped in the same commit.
Jul 28, 2026The first exact-page user-agent-labelled crawler request appeared. GPTBot- and ClaudeBot-labelled requests fetched the sitemap.
Jul 29, 2026A PerplexityBot-labelled request fetched the Research hub. Applebot- and PetalBot-labelled requests fetched the exact page.
Jul 31, 2026A GoogleOther-labelled request fetched the exact page. A PerplexityBot-labelled request fetched the sitemap.
Aug 1, 2026Perplexity cited the page for the target query. The citation repeated across three consecutive daily jobs and appeared in five regional observations.

In the window the page received nine successful main-page requests across seven user-agent-labelled crawler names: FacebookBot, Applebot, PetalBot, Bytespider, ChatGPT-User, GoogleOther and Meta-ExternalAgent. Some appeared more than once.

No record shows PerplexityBot or Perplexity-User fetching that exact page URL before the citation. PerplexityBot- labelled requests did reach the Research hub and the sitemap.

The page became discoverable and was later cited. The logs do not prove the route Perplexity used, and none of the user-agent-labelled events passed provider verification. The citation is reportable because it repeated across three consecutive daily jobs and five regional observations (repeated-run study). The full series, 16 frozen Perplexity observations with one later miss, is on How Stable Is a New AI Citation?

Timeline from the July 27 deployment through user-agent-labelled crawler requests on July 28, 29 and 31 to the first Perplexity citation on August 1, with the aggregate counts and the missing PerplexityBot exact-page record called out.
Figure 1. Observed sequence, not proven causality. Every crawler name shown is user-agent-labelled.

Can a page be cited without a visible crawler hit?

Yes. A missing exact-page log record does not mean the page was invisible. Possible explanations include:

  • the provider already held the page in an index or cache;
  • discovery happened through the Research hub, the sitemap, a link graph or a search partner;
  • the answer system used a conventional search index rather than a dedicated AI crawler;
  • the request came from infrastructure the classifier did not recognise;
  • logs were sampled, expired, filtered or stored in a different layer;
  • the cited source was selected from a search result without a new fetch reaching the origin.

A sitemap is a discovery aid, not a guarantee of crawling, indexing, retrieval or citation.

What the crawler log should retain

FieldWhy it matters
Timestamp and timezoneReconstruct the sequence and match it to prompt runs.
Source IP and verification resultSeparate verified providers from spoofed labels.
Full user-agentPreserve the original signal so events can be reclassified later.
Host, path and queryKnow exactly what was requested; redact sensitive values in reporting.
Status and content typeDistinguish successful HTML from redirects, errors and assets.
Bytes and latencySpot empty responses, truncation and operational failures.
Referer and cache indicatorsUnderstand routing and whether the origin served the request.
robots.txt and WAF decisionExplain allow, block, challenge and rate-limit outcomes.
Parser and verifier versionMake historical classifications reproducible.

Collect this at the origin server, CDN or edge. Analytics platforms filter bot traffic out before you see it.

Set robots policy by purpose

Not one “AI bots” switch. Decide separately whether to allow:

  • search discovery crawlers that make pages eligible for search-backed answers;
  • user-triggered agents that fetch a page on someone's request;
  • training or model-improvement crawlers;
  • ordinary search crawlers that support existing search visibility.

robots.txt is a crawler preference, not access control. The user-triggered agents may not follow it. Use authentication, network controls and rate limits for anything sensitive. Google-Extended and Applebot-Extended are control tokens, not user agents that should appear in logs.

Use a five-layer measurement model

LayerQuestionRecommended output
1. AccessCould the provider fetch the page?Verified successful requests by URL and role.
2. DiscoveryDid the system learn that the page exists?Sitemap and hub fetches, index evidence, discovery timestamps.
3. RetrievalWas the page selected for a prompt or agent action?Answer-time fetches, traces or provider evidence.
4. AnswerWas the page or brand mentioned or cited?Saved answer, citation URL, prompt, engine, region and time.
5. OutcomeDid visibility create value?Qualified visits, conversions, pipeline or support deflection.

Layers 1 and 2 come from your logs. Layers 3 and 4 come from repeated prompt runs with stored answers and cited URLs; the rules are in the measurement methodology and the prompts are chosen with the prompt-selection method.

What not to report

  • “Perplexity crawled us 43 times” when the count comes only from a user-agent regex.
  • “The page is indexed” because a crawler fetched the sitemap.
  • “The model used our content” because the server returned 200.
  • “This crawl caused the citation” because it happened earlier in the timeline.
  • “Crawler traffic equals AI traffic” when bot requests and human referrals are blended.
  • One combined crawler score that mixes search, user retrieval, training and ordinary search.

A practical audit workflow

  1. Choose the priority pages and the exact non-branded questions they answer.
  2. Confirm each page is public, indexable, canonical and present in the hub and sitemap.
  3. Collect raw request evidence at the server or edge, before analytics filters remove bots.
  4. Classify each agent by operational role using current provider documentation.
  5. Verify source identity with official IP or DNS methods wherever one is published.
  6. Report successful access separately from discovery, retrieval, answer and outcome evidence.
  7. Run the same prompts repeatedly across engines and regions; save the full answer and the cited URLs.
  8. Review deltas at fixed checkpoints and fix the weakest layer first.

To test whether one published page moved AI visibility, this workflow is the access-layer half of the first-citation trial protocol. Fix access before rewriting the article. A page that returns 403 to a search crawler is failing an access test, not a content test.

Crawler audit worksheet

The worksheet mirrors this page: fields to retain per request, role and verification method per agent, the verified / failed / not-verifiable classification, the five layers, the four robots decisions and a policy review date. No customer data.

Frequently asked questions

No. A verified, successfully served request proves the provider could read that URL at that time. Indexing, retrieval for a later question and citation in an answer are three further events, each needing its own evidence.

How this page was made

Crawler roles, control tokens and verification methods come from the provider documentation linked above, re-read on August 3, 2026. The AirPulse example uses aggregates from first-party request and answer-monitoring data over a frozen July 27 to August 1, 2026 window. IP addresses, customer identifiers, query strings and raw answers are excluded; every crawler name is user-agent-labelled, not verified. Timing is sequence, not causation. This page carries a quarterly review date.

Sources

  1. OpenAI: crawlers and user agents (OAI-SearchBot, ChatGPT-User, GPTBot)
  2. OpenAI: publishers and developers FAQ
  3. OpenAI: OAI-SearchBot IP ranges
  4. Perplexity: PerplexityBot and Perplexity-User documentation
  5. Perplexity: PerplexityBot IP ranges
  6. Anthropic: web crawlers and how to block them
  7. Google: common crawlers and control tokens
  8. Google: overview of Google crawlers
  9. Google: verify Googlebot and other Google crawlers
  10. Apple: about Applebot and Applebot-Extended
  11. Microsoft: how to verify Bingbot
  12. AirPulse: how to prove a GEO change improved AI visibility
  13. AirPulse: AI-search visibility measurement methodology

See which verified AI agents visit your site.

AirPulse AI Traffic separates crawler roles, keeps the raw evidence, and connects access data to the pages and prompts you care about.