Skip to the methodology
The Share of Voice methodReviewed

Built to be
questioned.

Your report should hold up when someone asks,
“Where did that number come from?”

Defined questions. Preserved answers. Visible limits.Method / 01

A benchmark with
the workings attached.

We measure how often your brand appears, is recommended and is named first in a defined set of freshly collected model API answers. You get the counts, the comparisons and the original evidence.

The scope is one brand, a category, an audience and a global or country market, in English, German or French. You approve the questions after payment and before collection. Results describe that test, at that time. They do not estimate all AI conversations, search demand or market share.

01 / Before

Agree the test.

The question set is saved before any paid answers are collected.

02 / During

Keep the evidence.

Original answers, prompts, model IDs, collection times and returned sources stay attached to the report.

03 / After

Show the arithmetic.

Counts and denominators sit beside the percentages. Missing analysis stays visible.

AI for interpretation. Code for measurement.

We favour deterministic rules wherever they can do the job. Software builds the paid question panel, checks collection integrity, retrieves quotations and calculates the scores. These steps do not ask a model to improvise a number.

Where AI adds value

The providers generate the answers being measured. Models also interpret website context, classify recommendations, review ambiguous product identities and help turn evidence into specific action briefs.

Where rules take over

Validation, supported name matching, deduplication, denominators, percentages and ranking follow defined code. With the same saved inputs, accepted labels, identity decisions and code version, the arithmetic is repeatable.

Deterministic counting does not make an AI label infallible. Repeating collection or interpretation can change the evidence or its labels. We preserve the record, show coverage and link findings to original answers so you can inspect the result.

The useful question

“For the buying decisions we tested, who gets recommended instead of us—and what evidence can we act on?”

Control the question.
Then repeat it.

You supply two to five buyer needs and two to ten competitors. Every need is tested on the same five criteria: overall fit, value for money, quality, customer experience and ease of adoption.

Choose the number of buyer needs
624 answers52 distinct prompts · 156 answers per provider

Each need gets five criteria. One extra group tests the category without the audience or buyer needs. Fewer needs mean more repeats, not a lower price.

The full sampling table
Current paid-report design
Buyer needsDistinct promptsRepeats per modelTotal answers
2228704
3325640
4424672
5523624

We choose the smallest whole number of repeats that reaches at least 624 answers, with a minimum of three. This is a collection-design target, not a statistical-power calculation. “600+ answers” includes repeats across models; it does not mean 600 unique questions or independent buyers.

What the collection model actually sees

The category, audience, market, buyer need and buying criterion go into a fixed template. The target brand, website and competitor list are not supplied as separate collection context. Validation blocks recognized tracked names, URLs and selected instruction patterns in the brief.

Example · generated by the current question builder
Which project management software for small design studios in Switzerland would you consider for this buyer need: “Plan client projects”? Compare suitable options on overall fit and explain the tradeoffs. If no suitable option can be established, say so. Do not substitute a different product category.

A separate category-only group removes the audience and individual needs. It is included in the total and shown separately in the measurement record. It offers a broader comparison, not a causal test of bias.

How website suggestions affect the brief

The website-only preview can use a model to interpret public page text and suggest a category, audience, market and need, with deterministic extraction as a fallback. Separate website suggestions offered while editing the paid brief use deterministic extraction. These are descriptions of the seller’s site, not measured customer demand.

You review the paid brief and approve the full question set. Use real sales calls, customer interviews or support requests to choose needs. Brand-name checks cannot prevent a highly specific feature bundle from favouring one supplier. Review warnings flag selected promotional phrases, crowded descriptions and near-duplicate needs; they cannot detect every leading question.

Four providers.
A documented setup.

We call one configured model API per provider. Each receives the same approved prompts and number of repeats. These are fresh requests without a consumer account’s chat history or personalisation.

01

OpenAI

openai/gpt-5-mini
02

Gemini

google/gemini-3-flash
03

Claude

anthropic/claude-sonnet-4.6
04

Perplexity

perplexity/sonar

Current defaults. Your report’s stored model IDs and collection times are the record for that run. Provider aliases and search results can change over time.

Search, location and generation settings

OpenAI, Gemini and Claude have native web-search tools enabled; the model can decide whether to use them. Perplexity uses its search-enabled Sonar model. Calls run through Vercel AI Gateway. Search availability and behaviour differ by provider, so this compares the configured systems rather than isolating model capability.

Shared generation instruction
Answer in English in at most 300 words. Use at most one web search query if needed.

OpenAI uses low search context and low reasoning effort. Claude’s search tool is capped at one use. The one-query and 300-word instructions are requests, not a guarantee that every provider implements them identically. The output budget is 4,000 tokens; empty and length-truncated responses are rejected. We do not set a common temperature or random seed.

Technical failures can be retried with the saved prompt. Recovery retains successful answers rather than collecting them again.

The selected country is written into the prompt. “Global” uses worldwide wording. Neither setting samples users by country or verifies local IP delivery. Results are observations in your selected question language, with no regional audience weighting.

API answers and consumer apps are different surfaces.

This report does not reproduce the ChatGPT, Gemini, Claude or Perplexity apps, their default models, personalisation or interface-specific search. Provider names identify the systems tested; they imply no partnership or endorsement.

The denominator
changes the story.

AI share of voice is a brand’s share of the counted brand mentions in a defined set of AI answers. A brand counts at most once per answer, even if it appears several times. A model extraction step labels the answer; code validates its evidence and aggregates the counts.

Worked example · illustrative data

The same 12 mentions.
Three different measures.

Of 40 analyzed answers12 / 40 = 30%

Mention rate. How often the brand appeared in an answer.

Of 24 mentions across selected brands12 / 24 = 50%

Selected-brand share of voice. Depends on which competitors you chose.

Of 48 mentions across all identified options12 / 48 = 25%

Discovered-option share of voice. Includes relevant alternatives beyond your list.

Several brands can appear in one answer, so brand-answer mentions can exceed the number of answers. “All identified options” means those found by this analysis, not every supplier in the market.

Named
Answers containing an extracted mention of the brand ÷ analyzed answers. A mention can be positive, negative or incidental.
Recommended
Answers labelled “recommended” or “top pick” for the brand ÷ analyzed answers. Caveats, alternatives and advice against it remain separate labels.
Named first
Answers where the brand is first among the retained entities in the text ÷ analyzed answers. Text order alone does not establish preference.
Your site in sources
Successful answers that cite your host or a subdomain ÷ all successful answers. This metric does not require extraction to succeed.
Shareable rank
Rank by recommendation count within the chosen report scope. Ties share a rank (1, 2, 2, 4). Ranking cards require full collection and analysis, all four providers, at least 24 analyzed answers in the selected scope and at least one recommendation for the brand. A card scoped to one buyer need describes that subset, not an overall market rank. Selecting the strongest subset does not make it representative of the whole report.
Aliases, category membership and citations

Customer-supplied target aliases and supported domain-name spelling variants are grouped. A separate source-linked identity review can reconcile references to the same product or remove category/channel labels. Uncertain matches stay separate, and unavailable or partial review is disclosed. Original answers remain unchanged.

Discovered competitors must be classified as relevant options; integration partners and incidental third-party names are excluded. Target-brand mentions are retained even when incidental or outside the requested category, so “named” is broader than “recommended as a suitable option”. Labels can be wrong; inspect the source answer.

Cited sources are the links the provider ties to the answer: ChatGPT annotations, Claude quoted citations, numbered Perplexity references and Gemini grounding sources. Each engine’s search queries and search results are recorded separately. Reports collected before October 2026 list Claude’s and Perplexity’s search results as their sources. For some Gemini redirect links, a domain-shaped source title supplies an unverified host. This is labelled in the report. A listed source does not establish that a specific passage was cited, that the page is accurate, or that it caused the recommendation.

Checks you can inspect.

The measurement record shows what passed, what was assessed and what is missing. Each safeguard addresses a specific failure mode; none is a certificate of neutrality.

  1. Approved questions, preserved

    The paid panel is fixed before collection. Answers are collected afresh; favourable preview answers are not substituted. Questions are not rewritten in response to results.

  2. The actual requests are checked

    Every planned question, phrasing, provider and repeat must be present once. The integrity check also verifies that the stored prompt matches the approved question.

  3. Extraction is blind to the customer’s identity

    The extraction model receives the question, category and answer passages. It is not told which brand is the customer, the tracked competitor list or the answer’s provider. Names remain in the text: this is target-blind annotation, not anonymisation. Target matching happens afterwards in code.

  4. Quotations come from the saved answer

    Code retrieves source passages and checks names and quotations against the answer text. The model classifies; it does not invent quotation text. Grounding verifies that evidence exists, not that the classification or product claim is true.

  5. Unhelpful answers stay in the record

    The same extraction call assesses whether an answer is on topic, mixed, off topic or unclear. Analyzed off-topic and unclear answers remain in the denominator. Missing assessments are “not assessed”, never an automatic pass.

Validation status and completion threshold

Regression tests cover prompt leakage, request integrity, balanced collection, grounding, deduplication and denominator handling. Developer-labelled synthetic fixtures and a separate synthetic holdout exercise extraction and question fit. This is internal validation, not an independent audit or a published real-world accuracy estimate across industries.

The extraction model is OpenAI GPT-5 mini, including for OpenAI-generated answers. Hiding the provider label does not eliminate model-family effects, recognition from the text, or systematic classification errors. No routine human annotation or independent second extractor is promised.

A completed paid report requires every planned valid answer and grounded analysis for at least 95% of them. The 95% is a delivery threshold, not a confidence level. Failed analysis is excluded from derived brand denominators and coverage is disclosed; it never counts as absence. If collection or analysis falls short, the report goes to review rather than being released as complete. Shareable ranking cards apply the stricter 100% analysis requirement.

Repeatable procedure.
Variable answers.

Repeats show how much the result moved within this collection. They do not make this a probability sample of buyer demand. Related questions, shared search sources and model behaviour can make answers dependent.

What we show

Counts by provider and repeat. Mention-rate differences between the two phrasings, using matched question/repeat pairs where both answers were analyzed. A separate category-only view. Coverage alongside the result.

What that establishes

Observed variation in the tested setup. It does not establish a population margin of error, a probability that the brand will be recommended, or a bias-free result. We publish no unsupported confidence score.

Every model receives the same collection allocation. Overall brand rates pool analyzed answers; if some analysis is missing, providers can contribute slightly different counts. Failed extraction can be systematic, so excluding it can still affect the result. The blend is not weighted by provider market share, prompt popularity or your customers’ usage.

Comparing a later report

Keep the category, audience, market, needs, approved phrasings, tracked names and aliases consistent. Check the model IDs, collection and analysis versions, dates, source mix and coverage before interpreting a change.

Provider updates, live search and response variation can move results even when your website has not changed. A rerun can follow the same procedure without reproducing the same answers. A before/after difference alone cannot attribute improvement to a content edit, review campaign or other intervention.

A finding should
give you somewhere to start.

The paid report includes a prioritised work plan. Each task connects an observed gap to supporting answers, a suggested owner, indicative effort, completion criteria and a measurement baseline.

Observed gapInspect the evidenceDecide what to change

The focus is work marketers can own: clearer website content, useful editorial coverage, accurate external profiles and reviews, and correcting unsupported claims. An absent phrase does not prove an absent product capability. Product changes are not a conclusion to draw from this benchmark alone.

What the website and action review covers

We inspect a limited set of public HTML pages and attempt a robots.txt check. An optional deeper review uses sitemap discovery, additional pages from your site and selected cited external pages. Page, size and time limits apply, as do crawler exclusions; successful partial checks are retained. This is a targeted evidence review, not a complete site audit.

Where eligible evidence and budget allow, a bounded model pass drafts more specific briefs and a separate model call critiques them. Quotes must match inspected text and measurement baselines are preserved. If enrichment is unavailable or fails validation, the grounded checklist remains. This is an automated review, not a human consulting engagement.

The crawl does not render JavaScript, establish search-index inclusion, inspect every page or test access from a provider’s network. A missing page in this limited crawl is not proof that it does not exist. Source co-occurrence does not establish influence. Tasks are hypotheses to verify; they promise no lift in rankings, traffic or revenue.

The details travel
with the result.

The report retains the approved prompts, original answers and returned sources. Its measurement record shows collection integrity, analysis coverage, scope checks and wording differences. Model IDs and versions let you check which method produced the result.

What you can audit

The scope, sampling design, counting rules, model setup and limitations are public. Your report adds the approved questions, collected answers and supporting evidence needed to challenge a finding. A reproducible calculation means applying the same rules to the same saved record; it does not promise identical answers from a new run.

Our source code, internal analysis prompts and operational tuning remain proprietary. They are not required reading to inspect a reported count. The method and evidence can be scrutinised without publishing the software that runs them.

Question protocol
buyer-panel-v2
Collection protocol
api-search-english-300w-v1
Analysis version
report-v2.1
Extractor
source-anchored-v4-multilingual
Collection allocation
Equal across providers and approved prompts
Unit of counting
One retained brand per answer, counted once
Research operator
Share of Voice by Serge · Superstellar LLC
Method reviewed
Free previews, older reports and corrections

A free preview tests one website-derived question across the four providers, without the paid panel’s breadth or repeated design. Eligible previews can be served from a cache for up to 24 hours. Treat the displayed collection time and scope as part of the result. A paid order collects a new panel after you approve the questions.

The public sample uses a fictional brand and illustrative answers. It demonstrates the report experience, not a live benchmark or validation result.

Older reports retain their saved questions and analysis. The previous combined-needs protocol and exact-name scoring are not directly comparable with the current panel. Older extractions are not retrospectively described as target-blind. Consult each report’s recorded method.

If a name has been misclassified or a number does not reconcile, contact hello@shareofvoice.app with the report and answer reference. We can investigate the evidence; we cannot certify the truth of every provider claim.

The thinking behind
the method.

These sources inform our choices about disclosure, prompt variation and model-based analysis. They do not certify Share of Voice or validate our commercial report.

  1. 01
    HELM: Holistic Evaluation of Language ModelsStanford CRFM · Liang et al., 2022

    A model-evaluation framework that makes scenarios, metrics, prompts and outputs inspectable. Relevant to our evidence-first reporting, not a validation of our scores.

  2. 02
    Standards for DisclosureAmerican Association for Public Opinion Research

    A useful disclosure reference for exact wording, sampling, weighting, processing and limits. Our benchmark is a model-output study, not a survey of people or an AAPOR-certified study.

  3. 03
    ProSA: Assessing and Understanding the Prompt Sensitivity of LLMsZhuo et al. · Findings of EMNLP, 2024

    Documents sensitivity to prompt wording across models and tasks. Our two phrasings expose some variation; they cannot cover every way a buyer might ask.

  4. 04
    Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-JudgeShi et al. · IJCNLP–AACL, 2025

    Studies position bias in model-based judging. Its setting differs from our extraction task; it supports treating model labels as fallible, not borrowing an accuracy claim.

Judge it for yourself

Open a report.
Follow the evidence.

Explore the illustrative sample, from the headline to the individual answers and action plan.

Inspect the sample report US$179 per paid report. Optional monthly reporting.