This protocol is designed for a narrow first-citation trial: one question neighbourhood, one primary intervention and a result that another person can audit. It does not promise a citation. It creates a fair way to tell whether the work improved measured AI visibility.
The short answer
To prove that a GEO or AEO change improved AI visibility, treat the change as an intervention in a time series. Your unit is one fixed prompt on one fixed engine observed repeatedly over time. Your primary outcomes are brand mention rate and own-domain citation rate. Third-party citations and source coverage are separate outcomes. The claim is stronger only when the target prompts improve more than matched control prompts under rules written before launch.
A citation after launch is evidence that a citation happened. It is not, by itself, evidence that your page caused it.
What this trial can and cannot answer
| Question | The trial can answer | The trial cannot prove |
|---|---|---|
| Did the target prompt-engine cells improve? | Yes, if the repeated target series moves more than controls under the registered rule. | It cannot prove the same effect for every prompt, market or engine. |
| Was the new page accessible and retrievable? | Yes, when status, rendering, indexing or retrieval checks and logs are retained. | Access alone does not guarantee answer selection or citation. |
| Did one change cause the result? | It can support a bounded causal interpretation when timing, controls and contamination are handled well. | It cannot isolate one change if several pages, links, prompts or systems changed at once. |
| Is the result permanent? | No. It describes the engines, prompts and window measured. | AI answer and retrieval systems can change after the trial. |
Why a one-run before-and-after test is weak
AI answers vary even when the wording appears unchanged. In AirPulse's frozen production study of 30,504 responses, consecutive mention status changed 9.7% of the time. Among prompt-engine cells observed at least three times, 35.9% contained both a mentioned and an unmentioned result during the window. Consecutive citation-source sets had a mean Jaccard overlap of 0.396.
Those figures do not mean that every prompt has a 9.7% chance of flipping or that every URL has a 60.4% chance of disappearing. They describe one production cohort. Their practical meaning is simpler: repeated observations are necessary, engines should be reported separately, and source movement should not be hidden inside a single score.
Evidence: the repeated-run study.
Definitions used in this protocol
| Term | Working definition |
|---|---|
| Prompt-engine cell | One exact prompt paired with one named answer engine or surface. |
| Valid run | A completed answer that passes the declared validation rules. Timeouts, refusals, empty answers and parser failures are excluded and reported separately. |
| Mention rate | Valid runs containing an approved brand mention divided by valid runs in the same prompt-engine-window scope. |
| Own-domain citation rate | Valid runs exposing at least one cited URL on the tracked brand's domain divided by valid runs. |
| Third-party citation rate | Valid runs citing an independent source that mentions, validates, compares or competes with the brand. |
| Source coverage | The distinct cited domains or URLs observed across valid runs. |
| Target prompt | A frozen question the intervention is designed to influence. |
| Control prompt | A matched question left untouched so background engine movement is visible. |
Step 0: confirm that measurement is possible
Do this before writing the page. A blocked or broken page turns a content trial into an access test. Check each relevant URL and keep evidence, not just screenshots.
- The URL returns a successful status without authentication, regional lock or a bot challenge.
- robots.txt and page-level directives allow the relevant search crawler where inclusion is desired.
- The important answer text is present in rendered output and has a stable URL.
- The canonical URL, noindex state, redirects and sitemap entry are correct.
- Server or CDN logs can distinguish verified crawler or fetcher traffic from a spoofed user agent where verification is available.
- The team records access, discovery or indexing, retrieval and final answer selection as separate states.
OpenAI says OAI-SearchBot access is important for inclusion in ChatGPT search summaries and citations. Perplexity separately documents PerplexityBot and Perplexity-User. Google documents Googlebot for Search crawling and advises verifying traffic by reverse DNS or published IP ranges. None of these sources promises a citation. See the OpenAI publishers and developers FAQ, Perplexity crawler documentation and Googlebot documentation.
Step 1: choose one question neighbourhood
A first trial should start with one specific buyer question and a small set of close reformulations. Broad category prompts are usually harder to interpret because they mix several buyer needs and a large competitive set. The goal is not to publish many thin pages. The goal is to answer one real question better than the available evidence does today.
Build 30 to 60 candidate questions from three evidence sources:
- Observed buyer language: sales calls, support questions, Search Console queries, community discussions and competitor-comparison questions.
- Available engine traces: generated search queries, retrieval results or grounding metadata when the surface exposes them.
- Controlled expansion: reformulations, related questions, implicit questions, comparisons, entity expansions and role- or context-specific variants.
Score each candidate against five gates:
| Gate | Pass condition |
|---|---|
| Buyer relevance | A real buyer asks the question while discovering, evaluating or validating a solution. |
| Answer advantage | AirPulse can answer with direct product knowledge, first-party data or a clearer method. |
| Competitive feasibility | The current result set is not permanently dominated by sources the trial cannot reasonably challenge. |
| Distribution fit | The answer can be supported on AirPulse and, where appropriate, by a useful public artifact or independent reference. |
| Measurement fit | The prompt has a clear brand-mention or citation outcome and a reasonable control set. |
A suitable AirPulse example is “Which AI crawlers are visiting my website?” because AgentPulse can use first-party request logs. It remains an example, not a predetermined winner. The final topic should come from the evidence collected for the trial.
Use the selection method: Which AI Search Prompts Should Your Brand Track?
Step 2: preregister the test before Day 0
Write the plan before publishing or materially changing the target page. Store it in a durable place. If a field changes later, create a new version and record why.
| Field | Required value |
|---|---|
| Decision question | The exact claim the trial is intended to support. |
| Target brand | Canonical name and approved aliases; include explicit false-positive exclusions. |
| Target prompts | Stable IDs, exact text, locale, intent, buyer stage and version. |
| Control prompts | Stable IDs and the matching logic used to select them. |
| Engines | Named engine and surface; record a model identifier when exposed. |
| Baseline | Start and end window, timezone, cadence and required valid runs. |
| Intervention | One primary page or change, its URL and the planned publication window. |
| Outcomes | Mention rate, own-domain citation rate, third-party citation rate and source stability kept separate. |
| Decision rule | What counts as positive, null or inconclusive. |
| Stop conditions | Access failure, inadequate valid runs, prompt drift, engine migration or contamination. |
| Evidence retention | Raw answer, sources, timestamps, available retrieval traces, parser version and validation flags. |
| Privacy rule | Aggregate-only public output unless a named customer has approved a case study. |
Step 3: freeze targets and matched controls
Use 8 to 12 target questions in the chosen neighbourhood. Select 8 to 12 controls with similar buyer stage, specificity and baseline visibility, but do not change the pages that primarily answer the controls during the trial.
A useful match checks:
- buyer stage and commercial intent;
- question specificity and answer format;
- baseline mention and citation level;
- competitive source mix;
- engine coverage and valid-run availability;
- seasonality or known campaign effects.
Controls do not need to be perfect twins. They need to make the major background changes visible. If a model or retrieval system changes and both targets and controls move together, a before-and-after difference on the target is less persuasive.
Step 4: collect the baseline
Start with at least seven valid runs for brand detection and eight for source coverage per prompt-engine cell, spread across two to four weeks. These are practical starting floors from external repeated-measurement research, not universal guarantees. Increase the run count when the result is near a decision boundary, the query is commercially important or source coverage continues to grow.
| Store for every run | Why it matters |
|---|---|
| Prompt ID, exact text and roster version | Prevents silent question changes. |
| Engine, surface and available model identifier | Keeps unlike systems from being pooled. |
| Request and response timestamp | Preserves order and measurement window. |
| Raw answer and matched brand span | Allows manual audit of mention classification. |
| Citation URLs and normalised domains | Separates own-domain, independent and competitor sources. |
| Available search queries or retrieval results | Shows retrieval evidence without treating it as a final citation. |
| Parser version, failure state and reviewer overrides | Makes data cleaning reproducible. |
Calculation contract: the measurement methodology.
Step 5: build one useful intervention
The intervention should be one page that resolves the target question. It should not be a page-per-variation fan-out. The page must remain useful when no AI engine is involved.
- State the direct answer in the first 80 words.
- Use a title and headings that match the actual buyer question without repeating it unnaturally.
- Include the smallest table or diagram that makes the decision easier.
- Use first-party data only when the cohort, method and limits can be explained.
- Name the author and reviewer and link material claims to primary sources.
- Define important terms and separate observations from recommendations.
- State what the evidence does not prove.
- Provide a stable canonical URL, crawlable HTML and a clear link from the Research hub.
Do not add:
- unsupported claims that one markup type, word count or distribution channel guarantees citation;
- dozens of thin pages for small prompt variations;
- hidden text, keyword stuffing or content written only for a crawler;
- customer-level production data, tenant identifiers, private prompts or raw answers;
- several simultaneous site-wide changes that make the result impossible to attribute.
Step 6: distribute the evidence without fabricating authority
Publish the canonical page on airpulse.ai. Add a genuinely useful public artifact when the topic supports one: a CSV, a reproducible notebook, a checklist or a GitHub example. Seek independent references through ordinary editorial work. A disclosed community answer or a walkthrough video can help a real reader discover the work, but distribution is not proof that the intervention caused a citation.
Use one canonical source of truth. Other formats should link back to it and add value rather than restating the same paragraph across many domains.
Step 7: verify the intervention before measuring the result
Record the actual publication time. Confirm the final URL, response status, canonical, rendered answer text, sitemap entry and relevant access rules. Save a content hash or version identifier. Treat the first verified live state as Day 0.
A server log showing one verified fetch is stronger access evidence than a robots.txt screenshot. An indexed URL is stronger discovery evidence than a submitted sitemap. Neither is evidence of retrieval for the target question unless the engine exposes that trace.
Step 8: repeat the same schedule
| Checkpoint | Required action | Interpretation |
|---|---|---|
| Day 0 | Verify the published version, access rules and evidence hash. | The intervention exists and is eligible for observation. |
| Day 30 | Check repeated target and control outcomes; inspect retrieval and citation sources. | Early retrieval or mention movement may be visible; avoid declaring victory from one run. |
| Day 45 | Review whether any target cells show repeated mentions or citations and whether controls moved. | A first citation remains an observation until it repeats under the plan. |
| Day 60 | Apply the registered decision rule and publish the result. | Classify positive, null or inconclusive; preserve the full evidence table. |
Keep the engine-level series separate. Do not rewrite the prompt wording after launch. Store failures separately and do not count timeouts or refusals as “brand absent.”
Step 9: decide the outcome
| Outcome | Required interpretation | What to publish |
|---|---|---|
| Positive | Targets improve more than controls under the registered rule, with adequate valid runs and no material contamination. | Target and control time series, effect definition, uncertainty, raw counts and limits. |
| Null | The intervention was accessible or retrievable, but the registered target outcome did not materially improve. | The same evidence table, including where the chain stopped. |
| Inconclusive | Access, sample size, engine change, prompt drift or overlapping interventions prevent a decision. | The failure condition, available observations and a revised protocol. |
For the first AirPulse trial, use one explicit expansion rule and label it as a trial rule, not an industry law. A workable rule is: expand the question neighbourhood only after at least three target queries hold a 30% or higher registered success rate for two consecutive weeks, while the matched controls do not show the same movement.
Minimum public result table
| Field | Report |
|---|---|
| Prompt | Stable ID plus readable text or a declared aggregate group. |
| Engine | Separate row per engine or a clearly nested engine view. |
| Window | Start, end and timezone. |
| Valid runs | Completed valid runs over planned valid runs; excluded runs separately. |
| Mention rate | Numerator, denominator, percentage and interval. |
| Own-domain citation rate | Numerator, denominator, percentage and interval. |
| Third-party citation rate | Numerator, denominator and percentage. |
| Source stability | Coverage plus comparable-set Jaccard summary. |
| Access state | Allowed, fetched, indexed or retrievable, or unknown. |
| Version | Prompt roster, parser, page intervention and measurement version. |
Common failure modes
| Failure | Why it breaks the claim | Correction |
|---|---|---|
| One screenshot before and one after | Normal answer variability is mistaken for improvement. | Use repeated prompt-engine observations across time. |
| Prompt wording changed | The unit of measurement changed. | Freeze text and version any addition. |
| All engines blended | Different engine behaviour is hidden. | Report each engine first. |
| No controls | Background retrieval or model changes look like an intervention effect. | Use matched untouched prompts. |
| Several changes launched together | The primary intervention cannot be isolated. | Declare one primary change or call the result a bundle test. |
| Failed calls counted as absence | Infrastructure errors depress mention rates. | Exclude and report invalid runs. |
| Access treated as citation | Early-stage evidence is promoted into a final outcome. | Keep the four doors separate. |
| Only a positive result is publishable | The analysis becomes outcome-dependent. | Publish null and inconclusive results too. |
How long and how much work
Plan for about eight calendar weeks and roughly three person-weeks of distributed effort for the first trial. The calendar time is longer than the labour time because the baseline and follow-up runs must be spread across time.
| Workstream | Typical effort |
|---|---|
| Question research and preregistration | 3–4 person-days |
| Baseline collection and quality checks | 3–4 person-days distributed across 2–4 weeks |
| Page, evidence artifact and technical implementation | 5–7 person-days |
| Distribution and access verification | 2–3 person-days |
| Follow-up analysis and public result | 3–4 person-days |
A small trial usually needs one substantive canonical page, one supporting public artifact where useful, a limited set of legitimate distribution actions and one published result. It does not require dozens of pages.
Preregistration worksheet
The downloadable worksheet mirrors the Step 2 table: prompt IDs, engines, controls, outcomes, decision rules, stop conditions and the privacy review. It contains no customer or tenant data.
