This protocol covers a narrow first-citation trial: one question neighbourhood, one primary intervention and a result another person can audit. It does not promise a citation. The first registered run has a published result, including one later miss, on How Stable Is a New AI Citation? The Day 60 verdict will follow the decision rule registered here.
Treat the change as an intervention in a time series. The unit is one fixed prompt on one fixed engine observed repeatedly. Primary outcomes are brand mention rate and own-domain citation rate; third-party citations and source coverage are separate.
The claim holds only when target prompts improve more than matched controls under rules written before launch. A citation after launch shows a citation happened, not that your page caused it.
What this trial can and cannot answer
| Question | The trial can answer | The trial cannot prove |
|---|---|---|
| Did the target prompt-engine cells improve? | Yes, if the repeated target series moves more than controls under the registered rule. | It cannot prove the same effect for every prompt, market or engine. |
| Was the new page accessible and retrievable? | Yes, when status, rendering, indexing or retrieval checks and logs are retained. | Access alone does not guarantee answer selection or citation. |
| Did one change cause the result? | It can support a bounded causal interpretation when timing, controls and contamination are handled well. | It cannot isolate one change if several pages, links, prompts or systems changed at once. |
| Is the result permanent? | No. It describes the engines, prompts and window measured. | AI answer and retrieval systems can change after the trial. |
Why a one-run before-and-after test is weak
AI answers vary even when the wording is unchanged. In AirPulse's frozen production study of 30,504 responses, consecutive mention status changed 9.7% of the time. Among prompt-engine cells observed at least three times, 35.9% contained both a mentioned and an unmentioned result. Consecutive citation-source sets had a mean Jaccard overlap of 0.396.
Those figures describe one cohort, not a 9.7% flip chance per prompt or a 60.4% disappearance chance per URL. They mean repeated observations are necessary and engines are reported separately. Evidence: the repeated-run study.
Definitions used in this protocol
| Term | Working definition |
|---|---|
| Prompt-engine cell | One exact prompt paired with one named answer engine or surface. |
| Valid run | A completed answer that passes the declared validation rules. Timeouts, refusals, empty answers and parser failures are excluded and reported separately. |
| Mention rate | Valid runs containing an approved brand mention divided by valid runs in the same prompt-engine-window scope. |
| Own-domain citation rate | Valid runs exposing at least one cited URL on the tracked brand's domain divided by valid runs. |
| Third-party citation rate | Valid runs citing an independent source that mentions, validates, compares or competes with the brand. |
| Source coverage | The distinct cited domains or URLs observed across valid runs. |
| Target prompt | A frozen question the intervention is designed to influence. |
| Control prompt | A matched question left untouched so background engine movement is visible. |
Step 0: confirm that measurement is possible
Do this before writing the page. A blocked page turns a content trial into an access test. Check each URL and keep evidence, not screenshots.
- The URL returns a successful status without authentication, regional lock or a bot challenge.
- robots.txt and page-level directives allow the relevant search crawler where inclusion is desired.
- The important answer text is present in rendered output and has a stable URL.
- The canonical URL, noindex state, redirects and sitemap entry are correct.
- Server or CDN logs can distinguish verified crawler or fetcher traffic from a spoofed user agent where verification is available.
- The team records access, discovery or indexing, retrieval and final answer selection as separate states.
OpenAI says OAI-SearchBot access matters for inclusion in ChatGPT search citations (publisher FAQ). Perplexity documents PerplexityBot and Perplexity-User separately (crawler documentation). Google advises verifying Googlebot by reverse DNS or published IP ranges (Googlebot documentation). None promises a citation. Verification methods: which AI crawlers visit your website.
Step 1: choose one question neighbourhood
Start with one specific buyer question and a small set of close reformulations. Broad category prompts mix several buyer needs and a large competitive set. Answer one real question better than the available evidence does.
Build 30 to 60 candidate questions from three sources:
- Observed buyer language: sales calls, support questions, Search Console queries, community discussions, comparison questions.
- Engine traces: generated search queries, retrieval results or grounding metadata where exposed.
- Controlled expansion: reformulations, related and implicit questions, comparisons, entity and role-specific variants.
Score each candidate against five gates:
| Gate | Pass condition |
|---|---|
| Buyer relevance | A real buyer asks the question while discovering, evaluating or validating a solution. |
| Answer advantage | AirPulse can answer with direct product knowledge, first-party data or a clearer method. |
| Competitive feasibility | The current result set is not permanently dominated by sources the trial cannot reasonably challenge. |
| Distribution fit | The answer can be supported on AirPulse and, where appropriate, by a useful public artifact or independent reference. |
| Measurement fit | The prompt has a clear brand-mention or citation outcome and a reasonable control set. |
One AirPulse example is “Which AI crawlers are visiting my website?”, answered in the crawler reference from first-party request logs. Selection method: Which AI Search Prompts Should Your Brand Track?
Step 2: preregister the test before Day 0
Write the plan before publishing or changing the target page. If a field changes later, create a new version and record why.
| Field | Required value |
|---|---|
| Decision question | The exact claim the trial is intended to support. |
| Target brand | Canonical name and approved aliases; include explicit false-positive exclusions. |
| Target prompts | Stable IDs, exact text, locale, intent, buyer stage and version. |
| Control prompts | Stable IDs and the matching logic used to select them. |
| Engines | Named engine and surface; record a model identifier when exposed. |
| Baseline | Start and end window, timezone, cadence and required valid runs. |
| Intervention | One primary page or change, its URL and the planned publication window. |
| Outcomes | Mention rate, own-domain citation rate, third-party citation rate and source stability kept separate. |
| Decision rule | What counts as positive, null or inconclusive. |
| Stop conditions | Access failure, inadequate valid runs, prompt drift, engine migration or contamination. |
| Evidence retention | Raw answer, sources, timestamps, available retrieval traces, parser version and validation flags. |
| Privacy rule | Aggregate-only public output unless a named customer has approved a case study. |
Step 3: freeze targets and matched controls
Use 8 to 12 target questions in the chosen neighbourhood and 8 to 12 controls with similar buyer stage, specificity and baseline visibility. Do not change the pages that answer the controls during the trial.
Match on:
- buyer stage and commercial intent;
- question specificity and answer format;
- baseline mention and citation level;
- competitive source mix;
- engine coverage and valid-run availability;
- seasonality or known campaign effects.
Controls need not be twins. They make background changes visible: if a model change moves targets and controls together, a before-and-after difference on the target is less persuasive.
Step 4: collect the baseline
Start with at least seven valid runs for brand detection and eight for source coverage per prompt-engine cell, spread across two to four weeks. These are practical floors from external research, not guarantees. Increase the count when the result is near a decision boundary or source coverage keeps growing.
| Store for every run | Why it matters |
|---|---|
| Prompt ID, exact text and roster version | Prevents silent question changes. |
| Engine, surface and available model identifier | Keeps unlike systems from being pooled. |
| Request and response timestamp | Preserves order and measurement window. |
| Raw answer and matched brand span | Allows manual audit of mention classification. |
| Citation URLs and normalised domains | Separates own-domain, independent and competitor sources. |
| Available search queries or retrieval results | Shows retrieval evidence without treating it as a final citation. |
| Parser version, failure state and reviewer overrides | Makes data cleaning reproducible. |
Calculation contract: the measurement methodology.
Step 5: build one useful intervention
One page that resolves the target question, not a page-per-variation fan-out. It must stay useful when no AI engine is involved.
- State the direct answer in the first 80 words.
- Match the title and headings to the buyer question without repeating it unnaturally.
- Include the smallest table or diagram that makes the decision easier.
- Use first-party data only when the cohort, method and limits can be explained.
- Name the author and link material claims to primary sources.
- Define terms and separate observations from recommendations.
- State what the evidence does not prove.
- Provide a stable canonical URL, crawlable HTML and a link from the Research hub.
Do not add:
- claims that one markup type, word count or channel guarantees citation;
- dozens of thin pages for small prompt variations;
- hidden text, keyword stuffing or content written only for a crawler;
- customer-level production data, tenant identifiers, private prompts or raw answers;
- several simultaneous site-wide changes that make the result impossible to attribute.
Step 6: distribute the evidence without fabricating authority
Publish the canonical page on airpulse.ai. Add a public artifact when the topic supports one (a CSV, a notebook, a checklist, a GitHub example). Seek independent references through ordinary editorial work. A disclosed community answer or walkthrough video can help a reader find the work, but distribution is not proof the intervention caused a citation. Other formats link back to the one canonical source.
Step 7: verify the intervention before measuring the result
Record the publication time. Confirm the final URL, response status, canonical, rendered answer text, sitemap entry and access rules. Save a content hash or version identifier. The first verified live state is Day 0.
A verified fetch in a server log beats a robots.txt screenshot as access evidence. An indexed URL beats a submitted sitemap as discovery evidence. Neither shows retrieval for the target question unless the engine exposes that trace.
Step 8: repeat the same schedule
| Checkpoint | Required action | Interpretation |
|---|---|---|
| Day 0 | Verify the published version, access rules and evidence hash. | The intervention exists and is eligible for observation. |
| Day 30 | Check repeated target and control outcomes; inspect retrieval and citation sources. | Early retrieval or mention movement may be visible; avoid declaring victory from one run. |
| Day 45 | Review whether any target cells show repeated mentions or citations and whether controls moved. | A first citation remains an observation until it repeats under the plan. |
| Day 60 | Apply the registered decision rule and publish the result. | Classify positive, null or inconclusive; preserve the full evidence table. |
Keep the engine-level series separate. Do not rewrite the prompt wording after launch. Store failures separately and do not count timeouts or refusals as “brand absent.”
Step 9: decide the outcome
| Outcome | Required interpretation | What to publish |
|---|---|---|
| Positive | Targets improve more than controls under the registered rule, with adequate valid runs and no material contamination. | Target and control time series, effect definition, uncertainty, raw counts and limits. |
| Null | The intervention was accessible or retrievable, but the registered target outcome did not materially improve. | The same evidence table, including where the chain stopped. |
| Inconclusive | Access, sample size, engine change, prompt drift or overlapping interventions prevent a decision. | The failure condition, available observations and a revised protocol. |
Use one explicit expansion rule and label it a trial rule. A workable one: expand the question neighbourhood only after at least three target queries hold a 30% or higher registered success rate for two consecutive weeks while the matched controls do not show the same movement.
Minimum public result table
| Field | Report |
|---|---|
| Prompt | Stable ID plus readable text or a declared aggregate group. |
| Engine | Separate row per engine or a clearly nested engine view. |
| Window | Start, end and timezone. |
| Valid runs | Completed valid runs over planned valid runs; excluded runs separately. |
| Mention rate | Numerator, denominator, percentage and interval. |
| Own-domain citation rate | Numerator, denominator, percentage and interval. |
| Third-party citation rate | Numerator, denominator and percentage. |
| Source stability | Coverage plus comparable-set Jaccard summary. |
| Access state | Allowed, fetched, indexed or retrievable, or unknown. |
| Version | Prompt roster, parser, page intervention and measurement version. |
Common failure modes
| Failure | Why it breaks the claim | Correction |
|---|---|---|
| One screenshot before and one after | Normal answer variability is mistaken for improvement. | Use repeated prompt-engine observations across time. |
| Prompt wording changed | The unit of measurement changed. | Freeze text and version any addition. |
| All engines blended | Different engine behaviour is hidden. | Report each engine first. |
| No controls | Background retrieval or model changes look like an intervention effect. | Use matched untouched prompts. |
| Several changes launched together | The primary intervention cannot be isolated. | Declare one primary change or call the result a bundle test. |
| Failed calls counted as absence | Infrastructure errors depress mention rates. | Exclude and report invalid runs. |
| Access treated as citation | Early-stage evidence is promoted into a final outcome. | Keep the four doors separate. |
| Only a positive result is publishable | The analysis becomes outcome-dependent. | Publish null and inconclusive results too. |
How long and how much work
About eight calendar weeks and roughly three person-weeks for the first trial. Calendar time exceeds labour time because baseline and follow-up runs are spread across weeks.
| Workstream | Typical effort |
|---|---|
| Question research and preregistration | 3-4 person-days |
| Baseline collection and quality checks | 3-4 person-days distributed across 2-4 weeks |
| Page, evidence artifact and technical implementation | 5-7 person-days |
| Distribution and access verification | 2-3 person-days |
| Follow-up analysis and public result | 3-4 person-days |
A small trial needs one canonical page, one public artifact where useful, a few legitimate distribution actions and one published result.
Preregistration worksheet
The worksheet mirrors the Step 2 table. No customer or tenant data.
