These run counts are a practical floor from external repeated-measurement research, not a guarantee of a fixed margin of error in every market.
AirPulse's frozen production study covered 30,504 prompt-engine responses. Mention status changed between consecutive observations 9.7% of the time. 35.9% of cells observed at least three times changed from mentioned to not mentioned, or the reverse, at least once during the 30-day window. Consecutive citation-source sets had a mean Jaccard overlap of 0.396. One answer is useful evidence of what happened once. It is not a reliable estimate of how visible a brand is.
The evidence in one table
| Measure | Frozen study | Fresh 30-day validation | What it means |
|---|---|---|---|
| Valid response observations | 30,504 | 30,146 | Prompt-engine answers included after the stated filters |
| Completed jobs | 456 | 469 | Production analysis jobs contributing responses |
| Cells observed at least three times | 1,334 | 1,507 | Brand × prompt × engine units used for stability classification |
| Cells that changed state at least once | 35.9% | 33.9% | A substantial share was neither always present nor always absent |
| Consecutive mention-state flip rate | 9.7% | 10.0% | The next valid observation disagreed with the prior one about one time in ten |
| Mean source-set Jaccard overlap | 0.396 | 0.390 | Consecutive cited-domain sets shared less than half their combined sources |
The fresh validation was calculated from the production read replica on July 20, 2026. It is not pooled into the frozen study. Keeping the cohorts separate prevents a daily moving window from silently changing the original headline.
Why one run gives the wrong level of confidence
An AI answer is generated, not retrieved as a fixed ranked list. The model, search layer, source set, prompt wording, account state and time can all affect what appears. A brand may be present in one answer and absent in the next even when the user repeats the same wording.
That does not make the answers useless. It changes the unit of measurement. The useful question is not “Did the brand appear?” It is “In how many valid runs did the brand appear, for which prompt, on which engine, over what period?”
What AirPulse measured
The fundamental unit was a prompt-engine cell: one fixed prompt paired with one fixed answer engine. Every valid run stored a timestamped answer and its available citation data.
For the frozen 30-day cohort, AirPulse analysed:
- 30,504 valid response observations;
- 456 completed production jobs;
- 1,559 distinct brand-prompt-engine cells;
- 1,334 cells with three or more observations;
- four reported engine groups: ChatGPT, Gemini, Google AI and Perplexity.
The mention calculation treated each valid response as either mentioning the tracked brand or not mentioning it. Consecutive observations were ordered by timestamp within the same cell. A flip occurred when the Boolean state changed.
The citation-source calculation converted valid source URLs to normalised domains. For each pair of consecutive source sets, Jaccard similarity was calculated as the size of the intersection divided by the size of the union. A score of 1 means the sets were identical. A score of 0 means they shared no domain.
Finding 1: mention status changed about one time in ten
Across all comparable consecutive observations, mention state flipped 9.7% of the time. The engine-level rates were different:
| Engine | Consecutive mention flip rate |
|---|---|
| ChatGPT | 8.5% |
| Gemini | 6.2% |
| Google AI | 14.7% |
| Perplexity | 9.3% |
The rates do not establish a permanent ranking of engine stability. They describe this production cohort, its prompt roster, its dates and AirPulse's parsing rules. They do show why engines should not be pooled into one opaque score.
Finding 2: more than one-third of repeated cells changed at least once
Among cells observed at least three times, 35.9%contained both a mentioned and an unmentioned result during the window. A single observation would misclassify some of these cells as “always visible” or “never visible,” depending on which day was chosen.
This metric is cell-level, not response-level. It does not mean that 35.9% of every prompt will flip next time. It means 35.9% of the sufficiently observed brand-prompt-engine units were mixed during the studied period.
Finding 3: cited sources moved more than brand mentions
Consecutive source-domain sets had a mean Jaccard overlap of 0.396 and a median of 0.333. The complement—about 60% to 67% of the combined set—was not shared. That is a set comparison, not a URL survival probability.
This distinction matters operationally. A brand mention can remain present while the pages supporting the answer change. Teams should therefore track at least three separate outcomes:
- brand mention rate;
- own-domain citation rate;
- third-party citation rate and source coverage.
AirPulse exposes these as separate views through Prompt Visibility and Citation Visibility. They should remain separate in exports and executive reports.
How many runs should a team start with?
Use seven brand-detection runs and eight source-coverage runs per prompt-engine cell as a starting floor. The recommendation comes from the paper Don't Measure Once: Measuring Visibility in AI Search. AirPulse's production study supports the direction—repeat measurements—but does not independently prove that seven and eight are universally sufficient.
Spread the runs across two to four weeks. Running eight times in ten minutes measures short-term generation variability. It does not capture day-to-day retrieval, model or index changes.
Increase the run count when:
- the estimated rate is near a decision boundary;
- the prompt is commercially important;
- the observed outcomes differ sharply by engine;
- source coverage is still growing with each run;
- a model, prompt roster or measurement method changed during the window.
A minimum reporting template
Every reported result should state:
- the exact prompt or a stable prompt identifier;
- the engine and relevant model/surface identifier;
- the valid-run numerator and denominator;
- the measurement dates and timezone;
- mention rate, own-domain citation rate and third-party citation rate separately;
- a confidence interval or equivalent uncertainty display;
- the prompt-roster and parser version;
- exclusions, failures and missing-source handling.
The full calculation and storage contract is in the AirPulse AI-search visibility measurement methodology. The prompt-roster design is in Which AI search prompts should your brand track?
What this study does not prove
This is an observational production cohort. It is not a random sample of all brands, prompts or AI users. It does not estimate the causal effect of publishing one page, adding schema or changing crawler access. Model and retrieval systems can change after the study period.
The study also does not support these shortcuts:
- “A brand has a 9.7% chance of flipping next time.”
- “Every URL has a 60.4% chance of disappearing.”
- “Seven runs always provide the same confidence.”
- “One blended score can compare all engines safely.”
- “Being indexed guarantees retrieval or citation.”
Access remains a prerequisite. OpenAI says publishers should allow OAI-SearchBot if they want content considered for ChatGPT search summaries and citations, while Google separates crawling, indexing and serving. See the OpenAI publisher FAQ and Google crawling documentation. Neither source promises citation.
Download and reproduce
The downloadable files contain only aggregate statistics. They do not contain customer names, tenant IDs, raw prompts or raw answers. The measurement methodology documents the calculations and exclusions.
