Crawler reference

Which AI Crawlers Visit Your Website — and What Are They Doing?

A log line labelled GPTBot or PerplexityBot is not a citation, and until the source is verified it is not even proof that the request came from that provider. Classify each request by role, verify its origin against the provider's published method, confirm the page was actually served, and keep crawler evidence separate from answer evidence.

Harsh SongraReviewed by the AirPulse research and data team

AI companies reach a website for several different reasons. One bot builds a search index. Another fetches a page because a person asked a question. A third collects material for model improvement. Ordinary search crawlers also supply pages that later appear in an AI answer. Those are four different jobs, and a single “AI bot traffic” number cannot tell you which one you are looking at.

The short answer

Most crawler activity falls into four useful groups:

  • Search discovery: a crawler finds and indexes pages for a search product.
  • User-triggered retrieval: an agent fetches a page while answering a specific person.
  • Training or model improvement: a crawler gathers material under the provider's training controls.
  • Ordinary search and utility crawling: conventional indexes and platform services fetch the page and may indirectly support AI experiences.

A successful, verified visit proves that the provider could access that URL at that time. It does not prove that the URL was indexed, that it was retrieved for a later prompt, that it was cited in an answer, or that it produced a business result.

Four events teams often mix up

EventWhat it meansStrong evidenceWhat it does not prove
AccessA crawler successfully fetched the URL.Verified request, usable status, correct content.Indexing, retrieval or citation.
Discovery / indexingA system knows the page exists and may store it.Provider console, index evidence or repeated verified discovery.That the page will be used for a specific answer.
RetrievalThe page was selected or fetched for a query or agent action.Provider trace, verified user-triggered fetch or answer-time evidence.That the answer cited or relied on it.
Answer selectionThe answer mentioned or cited the page or brand.Saved answer, citation URL, prompt, engine, region and timestamp.That a prior crawl alone caused the selection.
Four stages: access, discovery, retrieval and answer selection, each with its own kind of evidence.
Figure 1. Access, discovery, retrieval and answer selection are separate events. Evidence for one does not automatically prove the next.

Crawler and control reference

Provider names and policies change without notice. Use this as an operational map, then verify against the linked provider documentation. This table was last checked against those sources on .

ProviderAgent or tokenOperational roleWhat it doesMeasurement note
OpenAIOAI-SearchBotSearch discoverySurfaces websites in ChatGPT's search results, summaries and links.Verify against OpenAI's published ranges.
OpenAIChatGPT-UserUser-triggered retrievalFetches a page because a person or a custom GPT asked for it.OpenAI states robots.txt rules may not apply. A fetch is not a visible citation.
OpenAIGPTBotTraining / improvementCrawls content that may be used to train foundation models.Keep entirely separate from search visibility.
PerplexityPerplexityBotSearch discoveryBuilds and refreshes Perplexity's search index.Perplexity states it is not used for foundation-model training. Verify against published ranges.
PerplexityPerplexity-UserUser-triggered retrievalFetches pages in response to user actions.Perplexity states these requests generally ignore robots.txt. Access is not citation.
AnthropicClaude-SearchBotSearch discoverySupports search quality and result discovery for Claude.The role is distinct from training collection.
AnthropicClaude-UserUser-triggered retrievalFetches pages when a user asks Claude to access them.Check the status and the exact path served.
AnthropicClaudeBotTraining / improvementCollects material for model development under Anthropic's controls.Apply a separate robots policy if the decision differs from search.
GoogleGooglebotSearch discoveryBuilds Google Search's index, which can supply search and AI experiences.Use Google's reverse-and-forward DNS method.
GoogleGoogle-ExtendedControl tokenControls specified Gemini training and grounding uses.Not a separate HTTP user agent. It should not appear as a requester.
AppleApplebotSearch / platform discoverySupports Apple search products and suggestions.Verify using Apple's documented DNS method.
AppleApplebot-ExtendedControl tokenControls specified generative-AI uses of Applebot-crawled content.It does not crawl as a separate bot.
MicrosoftBingbotSearch discoveryBuilds Bing's index, which can support Microsoft search experiences.Verify using Bing's documented method.

Primary documentation: OpenAI crawlers, Perplexity bots, Anthropic web crawlers, Google common crawlers, About Applebot.

Four lanes grouping crawler agents by role: search discovery, user-triggered retrieval, training, and ordinary search or utility crawling, with control tokens called out as not being crawlers.
Figure 2. Sort traffic by what the agent is for, not by which company sent it. One provider can run three agents with three different jobs.

Why a user-agent name is not proof

A user-agent header is text supplied by the requester. Anyone can copy a known crawler name into it. AirPulse has observed requests labelled PerplexityBot attempting paths such as environment and configuration files. Those requests are consistent with scanners spoofing a trusted name, not with a search crawler doing normal discovery.

So label unverified records as user-agent-labelled. Upgrade an event to verified only when the source passes the provider's documented verification method. Where a provider publishes no reliable method, the honest classification is not verifiable, rather than a guess at the likely vendor.

How to verify a crawler visit

  1. Keep the raw timestamp, timezone, source IP, full user-agent, host, path, status and response content type.
  2. Identify the provider's official crawler documentation and the role assigned to that agent.
  3. Verify the source against published IP ranges or the provider's reverse-and-forward DNS method. OpenAI publishes ranges for OAI-SearchBot, ChatGPT-User and GPTBot; Perplexity publishes them for PerplexityBot and Perplexity-User; Google and Apple document DNS verification instead.
  4. Confirm the exact URL returned a usable response. A redirect loop, block page, 403, empty shell or 5xx is not successful access.
  5. Classify the event as verified, failed verification, or not verifiable. Do not force every event into a provider bucket.
  6. Correlate verified access with sitemap discovery, index evidence, prompt runs, citations and outcomes — but do not claim causality without stronger evidence.
A four-test chain turning a user-agent label into a classified event: record the label, verify source identity by IP or DNS, confirm the exact URL was served, then classify as verified, failed verification or not verifiable.
Figure 3. Every arrow is a separate test. Failing one downgrades the event; it never deletes the record, and passing all four still only proves access.

What happened on AirPulse's first cited research page

AirPulse examined the research page How do I prove that a GEO or AEO change improved AI visibility? and the exact non-branded query with the same wording. What follows is an observed sequence in AirPulse data over a frozen window. It is not proof that any single crawl caused the citation.

DateObserved event
Jul 27, 2026The page went live. The Research hub link, XML sitemap entry and llms.txt entry shipped in the same commit.
Jul 28, 2026The first exact-page user-agent-labelled crawler request appeared. GPTBot- and ClaudeBot-labelled requests fetched the sitemap.
Jul 29, 2026A PerplexityBot-labelled request fetched the Research hub. Applebot- and PetalBot-labelled requests fetched the exact page.
Jul 31, 2026A GoogleOther-labelled request fetched the exact page. A PerplexityBot-labelled request fetched the sitemap.
Aug 1, 2026Perplexity cited the page for the target query. The citation repeated across three consecutive daily jobs and appeared in five regional observations.

The aggregate: during the observed window the page received nine successful main-page requests across seven user-agent-labelled crawler names — FacebookBot, Applebot, PetalBot, Bytespider, ChatGPT-User, GoogleOther and Meta-ExternalAgent. Some names appeared more than once.

The important gap: AirPulse did not find an aggregate record of PerplexityBot or Perplexity-User fetching that exact page URL before the citation. It did observe PerplexityBot-labelled requests to the Research hub and to the sitemap.

The safe conclusion: the page became discoverable and was later cited. The logs do not prove the route Perplexity used to find or retrieve it, and none of these user-agent-labelled events has passed provider verification. The citation is reportable because it repeated: three consecutive daily jobs and five regional observations rather than a single screenshot. That is the standard set out in the repeated-run study.

Timeline from the July 27 deployment through user-agent-labelled crawler requests on July 28, 29 and 31 to the first Perplexity citation on August 1, with the aggregate counts and the missing PerplexityBot exact-page record called out.
Figure 4. Observed sequence, not proven causality. Every crawler name shown is user-agent-labelled.

Can a page be cited without a visible crawler hit?

Yes. A missing exact-page log record does not mean the page was invisible. Possible explanations include:

  • the provider already held the page in an index or cache;
  • discovery happened through the Research hub, the sitemap, a link graph or a search partner;
  • the answer system used a conventional search index rather than a dedicated AI crawler;
  • the request came from infrastructure the classifier did not recognise;
  • logs were sampled, expired, filtered or stored in a different layer;
  • the cited source was selected from a search result without a new fetch reaching the origin.

A sitemap is a discovery aid, not a guarantee of crawling, indexing, retrieval or citation.

What the crawler log should retain

FieldWhy it matters
Timestamp and timezoneReconstruct the sequence and match it to prompt runs.
Source IP and verification resultSeparate verified providers from spoofed labels.
Full user-agentPreserve the original signal so events can be reclassified later.
Host, path and queryKnow exactly what was requested; redact sensitive values in reporting.
Status and content typeDistinguish successful HTML from redirects, errors and assets.
Bytes and latencySpot empty responses, truncation and operational failures.
Referer and cache indicatorsUnderstand routing and whether the origin served the request.
robots.txt and WAF decisionExplain allow, block, challenge and rate-limit outcomes.
Parser and verifier versionMake historical classifications reproducible.

Collect this at the origin server, CDN or edge. Analytics platforms filter bot traffic out before you see it, and bot traffic is what this audit is made of.

Set robots policy by purpose

Do not use one vague “AI bots” switch. Decide separately whether you want to allow:

  • search discovery crawlers that make pages eligible for search-backed answers;
  • user-triggered agents that fetch a page on someone's request;
  • training or model-improvement crawlers;
  • ordinary search crawlers that support existing search visibility.

robots.txt is a crawler preference, not an access-control system. Two of the agents above may not follow it at all, because the fetch is user-initiated. Use authentication, authorisation, network controls and rate limits for anything sensitive. Remember also that Google-Extended and Applebot-Extended are control tokens, not separate HTTP user agents that should appear as requests in server logs.

Use a five-layer measurement model

LayerQuestionRecommended output
1. AccessCould the provider fetch the page?Verified successful requests by URL and role.
2. DiscoveryDid the system learn that the page exists?Sitemap and hub fetches, index evidence, discovery timestamps.
3. RetrievalWas the page selected for a prompt or agent action?Answer-time fetches, traces or provider evidence.
4. AnswerWas the page or brand mentioned or cited?Saved answer, citation URL, prompt, engine, region and time.
5. OutcomeDid visibility create value?Qualified visits, conversions, pipeline or support deflection.

Layers 1 and 2 come from your own logs. Layers 3 and 4 come from repeated prompt runs with stored answers and cited URLs — the calculation rules are in the measurement methodology, and the prompts they run against are chosen using the prompt-selection method.

What not to report

  • “Perplexity crawled us 43 times” when the count comes only from a user-agent regex.
  • “The page is indexed” because a crawler fetched the sitemap.
  • “The model used our content” because the server returned 200.
  • “This crawl caused the citation” because it happened earlier in the timeline.
  • “Crawler traffic equals AI traffic” when bot requests and human referrals are blended.
  • One combined crawler score that mixes search, user retrieval, training and ordinary search.

A practical audit workflow

  1. Choose the priority pages and the exact non-branded questions they answer.
  2. Confirm each page is public, indexable, canonical and present in the hub and sitemap.
  3. Collect raw request evidence at the server or edge, before analytics filters remove bots.
  4. Classify each agent by operational role using current provider documentation.
  5. Verify source identity with official IP or DNS methods wherever one is published.
  6. Report successful access separately from discovery, retrieval, answer and outcome evidence.
  7. Run the same prompts repeatedly across engines and regions; save the full answer and the cited URLs.
  8. Review deltas at fixed checkpoints and fix the weakest layer first.

If the goal is to test whether one published page moved AI visibility rather than to inventory crawler traffic, this workflow is the access-layer half of the first-citation trial protocol. Fix technical access before rewriting the article. A page that returns 403 to a search crawler is failing an access test, not a content test.

Crawler audit worksheet

The downloadable worksheet mirrors this page: the fields to retain per request, the operational role and verification method per agent, the verified / failed / not-verifiable classification, the five measurement layers, the four separate robots decisions and a provider-policy review date. It contains no customer data.

Frequently asked questions

No. A verified, successfully served request proves the provider could read that URL at that time. Indexing, retrieval for a later question and citation in an answer are three further events, and each needs its own evidence.

No. A sitemap makes discovery easier. It does not guarantee crawling, indexing, retrieval or citation. Treat a sitemap fetch as discovery evidence at most, never as index evidence.

Yes, and AirPulse observed exactly that. The provider may already hold the page in an index or cache, may discover it through the hub, sitemap or link graph, may rely on a conventional search layer, or may fetch from infrastructure your classifier does not recognise. Absent log evidence is not evidence of absence.

No. Google documents Google-Extended as a control token for Gemini training and grounding uses, not as an HTTP user agent. If you see it as a requester in logs, treat the record as suspect.

No. Apple describes it as a control over specified generative-AI uses of content that Applebot has already crawled. It does not crawl as a separate bot.

Not on its own. A user-agent header is text the requester chooses. Verify the source IP or DNS with the provider's official method and keep the raw request so the event can be reclassified later.

Usually yes. OpenAI separates GPTBot from OAI-SearchBot, Anthropic separates ClaudeBot from Claude-SearchBot, and Google and Apple expose control tokens for AI uses. Check each provider's current documentation before editing robots.txt, because the agents and tokens change.

Not always. OpenAI states that robots.txt rules may not apply to ChatGPT-User because the fetch is user-initiated, and Perplexity states that Perplexity-User requests generally ignore robots.txt. Use authentication and network controls, not robots.txt, for anything that must not be fetched.

For operations: verified successful requests to priority pages, split by crawler role. For marketing: repeated mentions or citations for a fixed prompt set, recorded with engine, region, timestamp and cited URL. Never use crawler counts as a proxy for citations.

How this page was made

Crawler roles, control tokens and verification methods come from the official provider documentation linked above, re-read on August 3, 2026. The AirPulse example uses privacy-safe aggregates from first-party request and answer-monitoring data over a frozen July 27 to August 1, 2026 window: IP addresses, customer identifiers, query strings and raw answers are excluded, and every crawler name in it is user-agent-labelled rather than verified. Observed timing is reported as sequence, not as proof of causation. Harsh Songra drafted the page; the AirPulse research and data team reviewed the claims before publication. Provider policies change without notice, so this page carries a quarterly review date rather than a permanent claim.

Sources

  1. OpenAI: crawlers and user agents (OAI-SearchBot, ChatGPT-User, GPTBot)
  2. OpenAI: publishers and developers FAQ
  3. OpenAI: OAI-SearchBot IP ranges
  4. Perplexity: PerplexityBot and Perplexity-User documentation
  5. Perplexity: PerplexityBot IP ranges
  6. Anthropic: web crawlers and how to block them
  7. Google: common crawlers and control tokens
  8. Google: overview of Google crawlers
  9. Google: verify Googlebot and other Google crawlers
  10. Apple: about Applebot and Applebot-Extended
  11. Microsoft: how to verify Bingbot
  12. AirPulse: how to prove a GEO change improved AI visibility
  13. AirPulse: AI-search visibility measurement methodology

See which verified AI agents visit your site.

AirPulse AI Traffic separates crawler roles, keeps the raw evidence, and connects access data to the pages and prompts you care about.