The short answer
Treat an AI visibility score as the result of a specified sample. Preserve the questions, models, search settings and raw answers, repeat the measurement, and show counts by provider. A movement can reflect answer variability or a changed test setup; it does not automatically prove your marketing worked.
The dashboard is up three points. Everyone would quite like this to be good news.
Before assigning credit, there is a less glamorous question: would the number have moved if you had done absolutely nothing? A repeatable process gives you a chance to answer that. A single screenshot does not.
Repeat the test before explaining the result
| Run | Answers naming your brand | Mention rate |
|---|---|---|
| 1 | 11 / 50 | 22% |
| 2 | 14 / 50 | 28% |
| 3 | 12 / 50 | 24% |
The observed range is 22–28%. If next week returns 26%, calling it an improvement over the first run’s 22% would be premature. You have already seen that much movement inside the original test. Inspect which questions changed and repeat the later measurement.
Three repeats are useful evidence of variability, but they do not establish a reliable population confidence interval. Related prompts and repeated model answers are not automatically independent observations. More decimal places will not fix that study design.
Check that you are still running the same test
- Question wording and weighting: did the panel or the mix of buying intents change?
- Provider and model: was there a model update, fallback or a different endpoint?
- Search and context: was web search available, and did instructions or conversation history change?
- Completion: did one provider fail more often, changing the balance of the pooled score?
- Scoring: did the alias list, competitor set or recommendation classifier change?
A failed request must remain visible. If a provider returns only 60 of 100 planned answers, show 60 valid and 40 missing. Do not treat the failures as answers without a mention. Equally, do not present the smaller sample as though the full plan completed.
Save the raw answer, citations and run metadata together. When a score changes, that record lets you separate a newly recommended rival from a parsing error or a newly recognized alias.
Why your ChatGPT screen can disagree with an API report
A provider’s consumer application and an API experiment can have different model configurations, instructions, search behavior and personal context. OpenAI documents that ChatGPT memory can use past chats and other available context to personalize responses. An isolated API request does not automatically reproduce your signed-in conversation.
That is evidence that the settings can differ. It does not prove a fixed ranking difference for every brand, or tell us which direction a particular score will move. You need a matched experiment to make that narrower claim.
Choose the measurement surface for the question. A controlled API study makes its inputs easier to record. A consumer-interface check shows an answer from that interface under those account and session conditions. Neither is a census of everything everyone sees. Our methodology describes our API scope explicitly.
Keep the comparison honest
Make a small change log: content releases, listing updates, model changes, panel revisions and outages. Keep provider-level counts next to the overall result. An apparent overall gain can come from a shift toward the provider where you were already strongest.
When the measurement setup materially changes, start a new baseline or run the old and new setups side by side. Do not splice them into an uninterrupted trend line. When your website changes, use unchanged questions as a reference where practical, and still acknowledge external changes you cannot control.
Useful monitoring does not require certainty about every movement. It requires knowing which movements justify opening the evidence, and having the restraint to leave a small wobble alone.
Sources & context
This guide responds to a question raised in Do you track AI visibility in any serious way, or is it mostly guesswork? on Reddit. The discussion informed the topic; it is not evidence of the effectiveness of any product mentioned there. Platform documentation is linked alongside the claims it supports.
Examples and figures are illustrative unless explicitly stated otherwise. Written and published by Share of Voice, a Superstellar LLC product. We sell AI visibility reports; our commercial perspective is worth keeping in mind.
Spotted something we should correct? Let us know.