Research protocol

How to Prove a GEO Change Improved AI Visibility

Do not use one ChatGPT screenshot as proof. Freeze the exact prompts, engines, run schedule, matching rules and controls before you publish. Measure a baseline, make one declared change, repeat the same measurements, and compare the target questions with untouched controls. Report the outcome as positive, null or inconclusive.

Harsh SongraReviewed by the AirPulse research and data team

This protocol is designed for a narrow first-citation trial: one question neighbourhood, one primary intervention and a result that another person can audit. It does not promise a citation. It creates a fair way to tell whether the work improved measured AI visibility.

The short answer

To prove that a GEO or AEO change improved AI visibility, treat the change as an intervention in a time series. Your unit is one fixed prompt on one fixed engine observed repeatedly over time. Your primary outcomes are brand mention rate and own-domain citation rate. Third-party citations and source coverage are separate outcomes. The claim is stronger only when the target prompts improve more than matched control prompts under rules written before launch.

A citation after launch is evidence that a citation happened. It is not, by itself, evidence that your page caused it.

What this trial can and cannot answer

QuestionThe trial can answerThe trial cannot prove
Did the target prompt-engine cells improve?Yes, if the repeated target series moves more than controls under the registered rule.It cannot prove the same effect for every prompt, market or engine.
Was the new page accessible and retrievable?Yes, when status, rendering, indexing or retrieval checks and logs are retained.Access alone does not guarantee answer selection or citation.
Did one change cause the result?It can support a bounded causal interpretation when timing, controls and contamination are handled well.It cannot isolate one change if several pages, links, prompts or systems changed at once.
Is the result permanent?No. It describes the engines, prompts and window measured.AI answer and retrieval systems can change after the trial.
Flow from question selection and preregistration through baseline, intervention, verification, repeated measurement and verdict, with matched controls as an unchanged rail beneath the flow.
Figure 1. The trial sequence. The rules are frozen before the intervention, and every outcome remains publishable.

Why a one-run before-and-after test is weak

AI answers vary even when the wording appears unchanged. In AirPulse's frozen production study of 30,504 responses, consecutive mention status changed 9.7% of the time. Among prompt-engine cells observed at least three times, 35.9% contained both a mentioned and an unmentioned result during the window. Consecutive citation-source sets had a mean Jaccard overlap of 0.396.

Those figures do not mean that every prompt has a 9.7% chance of flipping or that every URL has a 60.4% chance of disappearing. They describe one production cohort. Their practical meaning is simpler: repeated observations are necessary, engines should be reported separately, and source movement should not be hidden inside a single score.

Evidence: the repeated-run study.

Definitions used in this protocol

TermWorking definition
Prompt-engine cellOne exact prompt paired with one named answer engine or surface.
Valid runA completed answer that passes the declared validation rules. Timeouts, refusals, empty answers and parser failures are excluded and reported separately.
Mention rateValid runs containing an approved brand mention divided by valid runs in the same prompt-engine-window scope.
Own-domain citation rateValid runs exposing at least one cited URL on the tracked brand's domain divided by valid runs.
Third-party citation rateValid runs citing an independent source that mentions, validates, compares or competes with the brand.
Source coverageThe distinct cited domains or URLs observed across valid runs.
Target promptA frozen question the intervention is designed to influence.
Control promptA matched question left untouched so background engine movement is visible.

Step 0: confirm that measurement is possible

Do this before writing the page. A blocked or broken page turns a content trial into an access test. Check each relevant URL and keep evidence, not just screenshots.

  • The URL returns a successful status without authentication, regional lock or a bot challenge.
  • robots.txt and page-level directives allow the relevant search crawler where inclusion is desired.
  • The important answer text is present in rendered output and has a stable URL.
  • The canonical URL, noindex state, redirects and sitemap entry are correct.
  • Server or CDN logs can distinguish verified crawler or fetcher traffic from a spoofed user agent where verification is available.
  • The team records access, discovery or indexing, retrieval and final answer selection as separate states.

OpenAI says OAI-SearchBot access is important for inclusion in ChatGPT search summaries and citations. Perplexity separately documents PerplexityBot and Perplexity-User. Google documents Googlebot for Search crawling and advises verifying traffic by reverse DNS or published IP ranges. None of these sources promises a citation. See the OpenAI publishers and developers FAQ, Perplexity crawler documentation and Googlebot documentation.

Four stages: access, discovery, retrieval and answer selection, each with its own kind of evidence.
Figure 2. Access, discovery, retrieval and answer selection are separate stages. Report each one separately.

Step 1: choose one question neighbourhood

A first trial should start with one specific buyer question and a small set of close reformulations. Broad category prompts are usually harder to interpret because they mix several buyer needs and a large competitive set. The goal is not to publish many thin pages. The goal is to answer one real question better than the available evidence does today.

Build 30 to 60 candidate questions from three evidence sources:

  • Observed buyer language: sales calls, support questions, Search Console queries, community discussions and competitor-comparison questions.
  • Available engine traces: generated search queries, retrieval results or grounding metadata when the surface exposes them.
  • Controlled expansion: reformulations, related questions, implicit questions, comparisons, entity expansions and role- or context-specific variants.

Score each candidate against five gates:

GatePass condition
Buyer relevanceA real buyer asks the question while discovering, evaluating or validating a solution.
Answer advantageAirPulse can answer with direct product knowledge, first-party data or a clearer method.
Competitive feasibilityThe current result set is not permanently dominated by sources the trial cannot reasonably challenge.
Distribution fitThe answer can be supported on AirPulse and, where appropriate, by a useful public artifact or independent reference.
Measurement fitThe prompt has a clear brand-mention or citation outcome and a reasonable control set.

A suitable AirPulse example is “Which AI crawlers are visiting my website?” because AgentPulse can use first-party request logs. It remains an example, not a predetermined winner. The final topic should come from the evidence collected for the trial.

Use the selection method: Which AI Search Prompts Should Your Brand Track?

Step 2: preregister the test before Day 0

Write the plan before publishing or materially changing the target page. Store it in a durable place. If a field changes later, create a new version and record why.

FieldRequired value
Decision questionThe exact claim the trial is intended to support.
Target brandCanonical name and approved aliases; include explicit false-positive exclusions.
Target promptsStable IDs, exact text, locale, intent, buyer stage and version.
Control promptsStable IDs and the matching logic used to select them.
EnginesNamed engine and surface; record a model identifier when exposed.
BaselineStart and end window, timezone, cadence and required valid runs.
InterventionOne primary page or change, its URL and the planned publication window.
OutcomesMention rate, own-domain citation rate, third-party citation rate and source stability kept separate.
Decision ruleWhat counts as positive, null or inconclusive.
Stop conditionsAccess failure, inadequate valid runs, prompt drift, engine migration or contamination.
Evidence retentionRaw answer, sources, timestamps, available retrieval traces, parser version and validation flags.
Privacy ruleAggregate-only public output unless a named customer has approved a case study.

Step 3: freeze targets and matched controls

Use 8 to 12 target questions in the chosen neighbourhood. Select 8 to 12 controls with similar buyer stage, specificity and baseline visibility, but do not change the pages that primarily answer the controls during the trial.

A useful match checks:

  • buyer stage and commercial intent;
  • question specificity and answer format;
  • baseline mention and citation level;
  • competitive source mix;
  • engine coverage and valid-run availability;
  • seasonality or known campaign effects.

Controls do not need to be perfect twins. They need to make the major background changes visible. If a model or retrieval system changes and both targets and controls move together, a before-and-after difference on the target is less persuasive.

Step 4: collect the baseline

Start with at least seven valid runs for brand detection and eight for source coverage per prompt-engine cell, spread across two to four weeks. These are practical starting floors from external repeated-measurement research, not universal guarantees. Increase the run count when the result is near a decision boundary, the query is commercially important or source coverage continues to grow.

Store for every runWhy it matters
Prompt ID, exact text and roster versionPrevents silent question changes.
Engine, surface and available model identifierKeeps unlike systems from being pooled.
Request and response timestampPreserves order and measurement window.
Raw answer and matched brand spanAllows manual audit of mention classification.
Citation URLs and normalised domainsSeparates own-domain, independent and competitor sources.
Available search queries or retrieval resultsShows retrieval evidence without treating it as a final citation.
Parser version, failure state and reviewer overridesMakes data cleaning reproducible.

Calculation contract: the measurement methodology.

Step 5: build one useful intervention

The intervention should be one page that resolves the target question. It should not be a page-per-variation fan-out. The page must remain useful when no AI engine is involved.

  • State the direct answer in the first 80 words.
  • Use a title and headings that match the actual buyer question without repeating it unnaturally.
  • Include the smallest table or diagram that makes the decision easier.
  • Use first-party data only when the cohort, method and limits can be explained.
  • Name the author and reviewer and link material claims to primary sources.
  • Define important terms and separate observations from recommendations.
  • State what the evidence does not prove.
  • Provide a stable canonical URL, crawlable HTML and a clear link from the Research hub.

Do not add:

  • unsupported claims that one markup type, word count or distribution channel guarantees citation;
  • dozens of thin pages for small prompt variations;
  • hidden text, keyword stuffing or content written only for a crawler;
  • customer-level production data, tenant identifiers, private prompts or raw answers;
  • several simultaneous site-wide changes that make the result impossible to attribute.

Step 6: distribute the evidence without fabricating authority

Publish the canonical page on airpulse.ai. Add a genuinely useful public artifact when the topic supports one: a CSV, a reproducible notebook, a checklist or a GitHub example. Seek independent references through ordinary editorial work. A disclosed community answer or a walkthrough video can help a real reader discover the work, but distribution is not proof that the intervention caused a citation.

Use one canonical source of truth. Other formats should link back to it and add value rather than restating the same paragraph across many domains.

Step 7: verify the intervention before measuring the result

Record the actual publication time. Confirm the final URL, response status, canonical, rendered answer text, sitemap entry and relevant access rules. Save a content hash or version identifier. Treat the first verified live state as Day 0.

A server log showing one verified fetch is stronger access evidence than a robots.txt screenshot. An indexed URL is stronger discovery evidence than a submitted sitemap. Neither is evidence of retrieval for the target question unless the engine exposes that trace.

Step 8: repeat the same schedule

CheckpointRequired actionInterpretation
Day 0Verify the published version, access rules and evidence hash.The intervention exists and is eligible for observation.
Day 30Check repeated target and control outcomes; inspect retrieval and citation sources.Early retrieval or mention movement may be visible; avoid declaring victory from one run.
Day 45Review whether any target cells show repeated mentions or citations and whether controls moved.A first citation remains an observation until it repeats under the plan.
Day 60Apply the registered decision rule and publish the result.Classify positive, null or inconclusive; preserve the full evidence table.

Keep the engine-level series separate. Do not rewrite the prompt wording after launch. Store failures separately and do not count timeouts or refusals as “brand absent.”

Step 9: decide the outcome

OutcomeRequired interpretationWhat to publish
PositiveTargets improve more than controls under the registered rule, with adequate valid runs and no material contamination.Target and control time series, effect definition, uncertainty, raw counts and limits.
NullThe intervention was accessible or retrievable, but the registered target outcome did not materially improve.The same evidence table, including where the chain stopped.
InconclusiveAccess, sample size, engine change, prompt drift or overlapping interventions prevent a decision.The failure condition, available observations and a revised protocol.
Illustrative mention-rate time series for target and control prompts before and after a Day 0 intervention line: targets rise after Day 0 while controls stay flat.
Figure 3. The comparison that carries the claim, drawn as an illustrative shape rather than measured data. Publish the raw numerators and denominators in an accessible table alongside any real chart.

For the first AirPulse trial, use one explicit expansion rule and label it as a trial rule, not an industry law. A workable rule is: expand the question neighbourhood only after at least three target queries hold a 30% or higher registered success rate for two consecutive weeks, while the matched controls do not show the same movement.

Minimum public result table

FieldReport
PromptStable ID plus readable text or a declared aggregate group.
EngineSeparate row per engine or a clearly nested engine view.
WindowStart, end and timezone.
Valid runsCompleted valid runs over planned valid runs; excluded runs separately.
Mention rateNumerator, denominator, percentage and interval.
Own-domain citation rateNumerator, denominator, percentage and interval.
Third-party citation rateNumerator, denominator and percentage.
Source stabilityCoverage plus comparable-set Jaccard summary.
Access stateAllowed, fetched, indexed or retrievable, or unknown.
VersionPrompt roster, parser, page intervention and measurement version.

Common failure modes

FailureWhy it breaks the claimCorrection
One screenshot before and one afterNormal answer variability is mistaken for improvement.Use repeated prompt-engine observations across time.
Prompt wording changedThe unit of measurement changed.Freeze text and version any addition.
All engines blendedDifferent engine behaviour is hidden.Report each engine first.
No controlsBackground retrieval or model changes look like an intervention effect.Use matched untouched prompts.
Several changes launched togetherThe primary intervention cannot be isolated.Declare one primary change or call the result a bundle test.
Failed calls counted as absenceInfrastructure errors depress mention rates.Exclude and report invalid runs.
Access treated as citationEarly-stage evidence is promoted into a final outcome.Keep the four doors separate.
Only a positive result is publishableThe analysis becomes outcome-dependent.Publish null and inconclusive results too.

How long and how much work

Plan for about eight calendar weeks and roughly three person-weeks of distributed effort for the first trial. The calendar time is longer than the labour time because the baseline and follow-up runs must be spread across time.

WorkstreamTypical effort
Question research and preregistration3–4 person-days
Baseline collection and quality checks3–4 person-days distributed across 2–4 weeks
Page, evidence artifact and technical implementation5–7 person-days
Distribution and access verification2–3 person-days
Follow-up analysis and public result3–4 person-days

A small trial usually needs one substantive canonical page, one supporting public artifact where useful, a limited set of legitimate distribution actions and one published result. It does not require dozens of pages.

Preregistration worksheet

The downloadable worksheet mirrors the Step 2 table: prompt IDs, engines, controls, outcomes, decision rules, stop conditions and the privacy review. It contains no customer or tenant data.

Frequently asked questions

It is enough to report a first observed citation if the evidence is retained. It is not enough to claim stable visibility or causal improvement. Report the exact prompt, engine, time, cited URL and whether it repeated. Keep the trial running.

Use the engines that matter to the decision and can be measured consistently. The same prompt on two engines is two cells. Do not hide missing engines in a blended score.

Only under a declared versioning rule. A material change starts a new intervention version. Emergency fixes can be made, but record them and decide whether the original comparison remains valid.

No. Access is an early gate. Discovery, retrieval and answer selection remain separate. Provider documentation explains how to permit access; it does not promise ranking, recommendation or citation.

Publish it. A null result narrows the explanation and improves the next test. Check which door failed, then preregister the next intervention instead of rewriting the original success rule.

How this protocol was made

This protocol combines the AirPulse measurement contract, the frozen 30,504-response stability study, the July first-citation trial plan and primary crawler documentation. The MyChild example that informed the plan was post-hoc, single-brand and descriptive: its strongest query neighbourhood performed much better than broad queries, but that observation does not prove that the same distribution or content changes will cause an AirPulse citation. Any public AirPulse result must use aggregate-only data unless a customer has explicitly approved a named case study. Harsh Songra drafted the page; the AirPulse research and data team reviewed the claims before publication.

Sources

  1. AirPulse AI-Search Visibility Measurement Methodology
  2. How Many Times Should You Run an AI Search Prompt? Evidence From 30,504 Responses
  3. Which AI Search Prompts Should Your Brand Track?
  4. OpenAI: Publishers and Developers FAQ
  5. Perplexity crawler documentation
  6. Google: Googlebot and crawler verification documentation

Measure a fixed prompt roster with Prompt Visibility.

AirPulse runs your frozen targets and controls on the same schedule across engines, so the trial evidence collects itself.