Prompt-selection methodology

Which AI Search Prompts Should Your Brand Track?

Track the questions real buyers ask before they know your brand, while they compare options, and when they need proof. Start with 12 to 20 prompts grouped by buyer stage and use case. Keep the wording stable for measurement. Do not build the roster from keyword volume or an LLM brainstorm alone.

Harsh SongraReviewed by the AirPulse research and data team

This page sits in the research cluster because the prompt roster is the input set for the measurement methodology: change the roster and the measured result changes with it. Prompt selection can change the result more than another dashboard feature. The same brand can be absent for broad category questions and visible for a narrow question it answers unusually well. AirPulse's MyChild field note is one example: in a 45-prompt, four-engine run, the brand appeared in 23 of 180 response cells, concentrated in open-source and developer-intent questions.

Funnel showing observed buyer evidence, question grouping, prioritisation and a frozen prompt roster.
Figure 1. Begin with observed buyer language, use generation to expand it, and freeze only the questions that pass review.

Start with the buyer decision, not the category label

“AI visibility” can mean marketing visibility or software observability. “GEO” can mean generative engine optimization or geospatial analysis. A broad category phrase may pull the wrong market into the answer.

Write the decision first. Examples:

  • choose a platform that tracks brand mentions and citations across ChatGPT, Gemini and Perplexity;
  • understand why a competitor is recommended but the brand is absent;
  • find a tool that connects AI visibility with Search Console and analytics;
  • prove whether a content change improved answer-engine outcomes.

Then write questions a buyer could naturally ask to make that decision.

Use five evidence sources

1. Sales and customer calls

Pull exact questions from discovery, objections, procurement and renewal calls. Record the role, buyer stage and product area. Do not publish customer wording that contains names or confidential context.

2. CRM and support notes

Look for repeated pains, alternatives, requirements and “why not” questions. These often produce more useful evaluation prompts than top-of-funnel keywords.

Use observed queries to identify existing language and pages with demand. Search data does not reveal every AI question, but it grounds the roster in real phrasing.

4. Product and comparison pages

Convert product capabilities into buyer questions. “Citation history” becomes “Which tools show how citation sources change over time?” A capability is not a prompt until it is written as a decision question.

5. LLM and search expansion

Use ChatGPT, Gemini or another model to expand gaps after the observed evidence is organised. Label generated prompts. Review them for ambiguity, duplicates and invented demand.

Google recommends original information, clear sourcing and useful first-hand analysis in its people-first content guidance. The same principle applies to a measurement roster: generation can assist, but it should not replace evidence from buyers and the product.

Organise prompts by buyer stage

StagePurposeExample
ProblemNames the pain without a solution categoryWhy does my AI visibility score change between runs?
CategoryLooks for an approach or class of productHow do brands measure visibility in AI answers?
EvaluationCompares capabilities or criteriaWhich tools repeat prompts and show citation history?
VendorNames the brand or a direct comparisonAirPulse vs HubSpot AEO
ProofAsks whether an intervention workedHow do I prove that a GEO change improved citations?

A healthy roster includes non-branded discovery and evaluation questions. Branded prompts are useful for reputation monitoring, but they cannot show whether the brand enters a shortlist before the buyer knows its name.

Add use-case and audience qualifiers only when they change the answer

Qualifiers improve relevance when they represent a real constraint: enterprise governance, B2B SaaS, healthcare compliance, local services or a specific integration. They create noise when they are inserted only to manufacture more pages.

Keep one broad parent question and a small number of decision-relevant variants. Do not create one prompt and one article for every wording change.

The MyChild case: one brand, three question neighbourhoods

AirPulse's corrected 29/100 field-note analysis shows why prompt fit matters. A later production run used 45 prompts across ChatGPT, Gemini, Google AI and Perplexity. The public aggregate download contains the grouped counts without customer identifiers, raw prompts or raw answers.

Question groupMention cellsRate within group
Explicit open-source prompts12/2842.9%
Adjacent developer prompts9/2437.5%
Other child-development prompts2/1281.6%
Bar chart comparing MyChild mention rates for open-source, developer-intent and other child-development prompts.
Figure 2. The association was large but post-hoc. It shows prompt-neighbourhood dependence; it does not prove that the words “open source” caused the result.

The open-source/developer groups were about 33.75 times more likely than the other group to produce at least one prompt-level mention in that run. The classification was created after observing the data, involved one brand and did not preserve generated fan-out queries. Treat it as a strong descriptive example, not a universal rule.

Score each candidate question before adding it

Use a simple 0–2 review for six criteria:

Criterion012
Buyer evidenceGenerated onlyIndirect signalSeen in calls, CRM or search data
Commercial intentInformational onlyAdjacent to decisionDirectly affects selection or action
Product fitAirPulse cannot helpPartial connectionDirect product workflow
Evidence advantageGeneric adviceSome internal examplesOriginal data or repeatable test
Wording clarityAmbiguous categoryNeeds qualifierClear question and market
Measurement valueNo stable outcomeSecondary diagnosticMention, citation or retrieval outcome

Prioritise prompts scoring 9 or more out of 12. Keep lower-scoring questions in a backlog until better evidence appears.

Build the first roster

For a first 12–20-prompt roster:

  1. choose three to five business decisions;
  2. add one problem question per decision;
  3. add one category or evaluation question per decision;
  4. add the strongest use-case qualifier where it changes the answer;
  5. include two or three branded/vendor prompts for reputation context;
  6. remove duplicates and ambiguous category phrases;
  7. assign intent, buyer stage, owner and success metric;
  8. freeze the wording and version the roster.

Do not optimise the roster after seeing the baseline. If a new wording is genuinely better, add it as a new version and preserve the old series.

Match every prompt to an evidence page

The roster is also a content map. Every important prompt should point to one canonical page that can answer it completely. Several related questions can belong to one page.

For this AirPulse cluster:

This cluster structure creates a clear internal-link path without publishing thin near-duplicates.

Review quarterly, measure continuously

Keep a frozen core roster for trend continuity. Review it quarterly, or when the product category, audience or buyer language changes materially.

During review:

  • retain high-value prompts even when the brand is absent;
  • retire prompts that no longer represent a buyer decision;
  • add new observed questions as a new roster version;
  • separate market changes from measurement changes;
  • record why each prompt was added, edited or removed.

Use the AirPulse measurement methodology for run counts, storage and reporting. Use Prompt Analysis to group and review the roster inside the product.

Download the worksheet

The worksheet is a CSV template with fields for prompt, buyer stage, evidence source, commercial intent, product fit, evidence advantage, wording clarity, measurement value, total score, owner and version.

Frequently asked questions

Start with 12 to 20 decision-relevant prompts. A smaller roster measured repeatedly is more useful than hundreds of generated prompts run once.

Include a few branded prompts for reputation and comparison monitoring, but keep most acquisition prompts non-branded so the study measures category discovery.

No. Group closely related questions into one complete page. Create a separate page only when the reader intent, answer or evidence is materially different.

Change it only when buyer language or the decision changes. Version the roster and do not splice the new wording into the old trend line.

It can propose candidates. A human owner should validate each against observed buyer evidence, product fit and a measurable outcome.

How this page was made

AirPulse analysed aggregate records from its production read replica in a read-only transaction. The public study excludes customer names, tenant identifiers, raw prompts and raw answers. The analysis code groups observations by brand, prompt and engine, orders them by time, and compares consecutive valid runs. Harsh Songra drafted the page; the AirPulse research and data team reviewed the claims and the downloadable aggregates before publication. AI assisted with structure and editing, but the published numbers must reproduce from the frozen evidence file.

References

  1. MyChild question-neighbourhood aggregate CSV
  2. AirPulse AI-search measurement research hub
  3. AirPulse AI-search visibility methodology
  4. Google: Creating helpful, reliable, people-first content

Build the roster inside the product.

Prompt Analysis groups candidate questions, tracks roster versions, and connects each prompt to its measured outcome.