Original study

How Many Times Should You Run an AI Search Prompt? Evidence From 30,504 Responses

Do not measure an AI-search prompt once. Start with at least seven runs per prompt and engine for brand-presence measurement and eight for source-coverage measurement, spread across two to four weeks, and report each engine separately.

Harsh SongraReviewed by the AirPulse research and data team

These run counts are a practical floor from external repeated-measurement research, not a guarantee of a fixed margin of error in every market.

AirPulse's frozen production study covered 30,504 prompt-engine responses. Mention status changed between consecutive observations 9.7% of the time. 35.9% of cells observed at least three times changed from mentioned to not mentioned, or the reverse, at least once during the 30-day window. Consecutive citation-source sets had a mean Jaccard overlap of 0.396. One answer is useful evidence of what happened once. It is not a reliable estimate of how visible a brand is.

Bar chart of AirPulse's frozen study findings: 9.7 percent consecutive mention flips, 35.9 percent mixed-state cells, and 60.4 percent source non-overlap.
Figure 1. Frozen 30-day production cohort. Source non-overlap is one minus the 0.396 mean Jaccard overlap; it is not the probability that one URL disappears.

The evidence in one table

MeasureFrozen studyFresh 30-day validationWhat it means
Valid response observations30,50430,146Prompt-engine answers included after the stated filters
Completed jobs456469Production analysis jobs contributing responses
Cells observed at least three times1,3341,507Brand × prompt × engine units used for stability classification
Cells that changed state at least once35.9%33.9%A substantial share was neither always present nor always absent
Consecutive mention-state flip rate9.7%10.0%The next valid observation disagreed with the prior one about one time in ten
Mean source-set Jaccard overlap0.3960.390Consecutive cited-domain sets shared less than half their combined sources

The fresh validation was calculated from the production read replica on July 20, 2026. It is not pooled into the frozen study. Keeping the cohorts separate prevents a daily moving window from silently changing the original headline.

Why one run gives the wrong level of confidence

An AI answer is generated, not retrieved as a fixed ranked list. The model, search layer, source set, prompt wording, account state and time can all affect what appears. A brand may be present in one answer and absent in the next even when the user repeats the same wording.

That does not make the answers useless. It changes the unit of measurement. The useful question is not “Did the brand appear?” It is “In how many valid runs did the brand appear, for which prompt, on which engine, over what period?”

Diagram comparing a misleading one-run screenshot with a repeated-run distribution across time.
Figure 2. A screenshot is an observation. A repeated series estimates a rate and shows its spread.

What AirPulse measured

The fundamental unit was a prompt-engine cell: one fixed prompt paired with one fixed answer engine. Every valid run stored a timestamped answer and its available citation data.

For the frozen 30-day cohort, AirPulse analysed:

  • 30,504 valid response observations;
  • 456 completed production jobs;
  • 1,559 distinct brand-prompt-engine cells;
  • 1,334 cells with three or more observations;
  • four reported engine groups: ChatGPT, Gemini, Google AI and Perplexity.

The mention calculation treated each valid response as either mentioning the tracked brand or not mentioning it. Consecutive observations were ordered by timestamp within the same cell. A flip occurred when the Boolean state changed.

The citation-source calculation converted valid source URLs to normalised domains. For each pair of consecutive source sets, Jaccard similarity was calculated as the size of the intersection divided by the size of the union. A score of 1 means the sets were identical. A score of 0 means they shared no domain.

Finding 1: mention status changed about one time in ten

Across all comparable consecutive observations, mention state flipped 9.7% of the time. The engine-level rates were different:

EngineConsecutive mention flip rate
ChatGPT8.5%
Gemini6.2%
Google AI14.7%
Perplexity9.3%

The rates do not establish a permanent ranking of engine stability. They describe this production cohort, its prompt roster, its dates and AirPulse's parsing rules. They do show why engines should not be pooled into one opaque score.

Finding 2: more than one-third of repeated cells changed at least once

Among cells observed at least three times, 35.9%contained both a mentioned and an unmentioned result during the window. A single observation would misclassify some of these cells as “always visible” or “never visible,” depending on which day was chosen.

This metric is cell-level, not response-level. It does not mean that 35.9% of every prompt will flip next time. It means 35.9% of the sufficiently observed brand-prompt-engine units were mixed during the studied period.

Finding 3: cited sources moved more than brand mentions

Consecutive source-domain sets had a mean Jaccard overlap of 0.396 and a median of 0.333. The complement—about 60% to 67% of the combined set—was not shared. That is a set comparison, not a URL survival probability.

This distinction matters operationally. A brand mention can remain present while the pages supporting the answer change. Teams should therefore track at least three separate outcomes:

  1. brand mention rate;
  2. own-domain citation rate;
  3. third-party citation rate and source coverage.

AirPulse exposes these as separate views through Prompt Visibility and Citation Visibility. They should remain separate in exports and executive reports.

How many runs should a team start with?

Use seven brand-detection runs and eight source-coverage runs per prompt-engine cell as a starting floor. The recommendation comes from the paper Don't Measure Once: Measuring Visibility in AI Search. AirPulse's production study supports the direction—repeat measurements—but does not independently prove that seven and eight are universally sufficient.

Spread the runs across two to four weeks. Running eight times in ten minutes measures short-term generation variability. It does not capture day-to-day retrieval, model or index changes.

Increase the run count when:

  • the estimated rate is near a decision boundary;
  • the prompt is commercially important;
  • the observed outcomes differ sharply by engine;
  • source coverage is still growing with each run;
  • a model, prompt roster or measurement method changed during the window.

A minimum reporting template

Every reported result should state:

  • the exact prompt or a stable prompt identifier;
  • the engine and relevant model/surface identifier;
  • the valid-run numerator and denominator;
  • the measurement dates and timezone;
  • mention rate, own-domain citation rate and third-party citation rate separately;
  • a confidence interval or equivalent uncertainty display;
  • the prompt-roster and parser version;
  • exclusions, failures and missing-source handling.

The full calculation and storage contract is in the AirPulse AI-search visibility measurement methodology. The prompt-roster design is in Which AI search prompts should your brand track?

What this study does not prove

This is an observational production cohort. It is not a random sample of all brands, prompts or AI users. It does not estimate the causal effect of publishing one page, adding schema or changing crawler access. Model and retrieval systems can change after the study period.

The study also does not support these shortcuts:

  • “A brand has a 9.7% chance of flipping next time.”
  • “Every URL has a 60.4% chance of disappearing.”
  • “Seven runs always provide the same confidence.”
  • “One blended score can compare all engines safely.”
  • “Being indexed guarantees retrieval or citation.”

Access remains a prerequisite. OpenAI says publishers should allow OAI-SearchBot if they want content considered for ChatGPT search summaries and citations, while Google separates crawling, indexing and serving. See the OpenAI publisher FAQ and Google crawling documentation. Neither source promises citation.

Download and reproduce

The downloadable files contain only aggregate statistics. They do not contain customer names, tenant IDs, raw prompts or raw answers. The measurement methodology documents the calculations and exclusions.

Frequently asked questions

No. It is enough to prove that one answer occurred. It is not enough to estimate a stable mention or citation rate.

Not in the primary report. Report ChatGPT, Gemini, Google AI and Perplexity separately because their generation, search and source behaviour differ. A blended summary may be added only when the engine mix and weighting are visible.

No. Spread the starting sample across two to four weeks. A short burst can be useful for debugging, but it does not replace a time-distributed measurement.

No. A mention means the answer names the brand. A citation means the answer links or attributes a source. A brand can be mentioned without its own site being cited, and its site can occasionally be cited without a clear brand mention.

No. Access may be necessary for direct crawling or user-requested fetching, but discovery, indexing, retrieval and final answer selection remain separate stages.

How this page was made

AirPulse analysed aggregate records from its production read replica in a read-only transaction. The public study excludes customer names, tenant identifiers, raw prompts and raw answers. The analysis code groups observations by brand, prompt and engine, orders them by time, and compares consecutive valid runs. Harsh Songra drafted the page; the AirPulse research and data team reviewed the claims and the downloadable aggregates before publication. AI assisted with structure and editing, but the published numbers must reproduce from the frozen evidence file.

Sources

  1. Don't Measure Once: Measuring Visibility in AI Search (arXiv)
  2. Quantifying Uncertainty in AI Visibility (arXiv)
  3. Google: Creating helpful, reliable, people-first content
  4. OpenAI: Publishers and Developers FAQ
  5. Perplexity crawler documentation
  6. AirPulse AI-search measurement research hub

See the repeated runs on your own prompts.

Prompt Visibility applies this protocol to your roster: repeated runs per engine, mentions and citations reported separately.