The unit of measurement
You cannot measure "ChatGPT" as a whole. You measure one fixed prompt on one engine, observed repeatedly. Everything else is arithmetic on top of that unit. A "visibility score" that does not tell you the prompts, the engine and the dates behind it cannot be reproduced, and a number that cannot be reproduced cannot show change.
Repetition matters because the same prompt does not return the same answer. In a frozen study of 30,504 production responses across four engines, the mention state of a prompt flipped between consecutive runs 9.7% of the time, and about a third of well-observed prompt-engine pairs showed both a mentioned and an unmentioned result within a month. One screenshot is one draw from that distribution.
Step 1: pick the prompts
- Eight to sixteen. Fewer and one flip moves the percentage by a lot; more and you will not keep the cadence.
- Write them as the buyer types them. "tool to see why users drop off during onboarding", not "product analytics platform". If sales hears the question on calls, it belongs in the set.
- Include two or three brand prompts ("what is [brand]", "[brand] vs [competitor]") so you can see the gap between being known and being recommended.
- Freeze the wording. Changing a prompt starts a new series; do not edit the old one in place.
Step 2: run them the same way each time
Use ChatGPT with search on, in a fresh conversation, with no memory of earlier prompts. Paste the full answer and its source list into a sheet with the date. If you are also tracking Perplexity, Gemini or Google AI Overviews, run those separately and keep them in separate columns. Never pool engines.
Step 3: score three things
| Column | Definition | Why it is separate |
|---|---|---|
| Named | Your brand appears in the answer text | This is what the reader acts on |
| Cited URL | Which of your URLs, if any, is linked as a source | Proves your page was retrieved; tells you which page is doing the work |
| Competitor named | Which other brands appear in the text | Shows who owns the job today and whether the shortlist is changing |
Then compute one figure per engine: visibility = prompts where you were named ÷ prompts run. Keep cited as its own figure. Do not add them. A brand can be named on 3 of 12 buying prompts and cited on 0 of 12, and that pair of numbers is the diagnosis: the model knows the name and has no page of yours for the work.
Step 4: repeat on a fixed cadence
Weekly is the minimum that produces a trend within a quarter. Daily is better if it is automated. What matters is that the interval is fixed and written down, so a change in the number is a change in the answers, not a change in when you looked.
| Week | Engine | Named | Cited | Reading |
|---|---|---|---|---|
| 1 | ChatGPT | 2 / 12 | 0 / 12 | Known on brand prompts, absent on the work |
| 2 | ChatGPT | 2 / 12 | 1 / 12 | One job page entered the retrieved set |
| 3 | ChatGPT | 4 / 12 | 2 / 12 | Named followed cited on the two job prompts |
What to leave out of the method
- A single blended score across engines and prompts. It hides the win and the zero in one number.
- A count of "engines covered". Report the engines you actually ran, by name.
- Sentiment and position rankings, at least at first. Named and cited are the two facts you can defend in a meeting.
- Referral traffic as the primary metric. Most of the reading happens inside the answer; see why AI search visits look empty in analytics.
