AI Share of Voice
Measure your brand visibility in AI-generated responses.
AI Share of Voice for the number you have to defend in the boardroom
You typed your brand and two competitors into ChatGPT, got a list that named them and not you, and now someone wants a number for Monday's deck. We run the same named-brand prompts across engines and hand back a percentage — but we also hand back the margin of error, because at the sample sizes most people run, that number moves around more than any campaign you'll launch this quarter. If you want to know exactly which sources are taking your share, that's a separate, sharper question.
Built and reviewed by Sayed Hasan, founder of SEOs Hut · Updated 28 September 2026 · Free, no signup
How many prompts do I need to track for reliable AI SoV?
More than most trackers run. At 10 prompts the worst-case 95% margin of error is about ±31 percentage points; at 30 prompts it's about ±18 points; you need roughly 385 runs to reach ±5 points. Below about 30–50 runs, a week-on-week change smaller than the margin of error is statistical noise, not a real shift — check the table on this page before you report a move.
That percentage is an estimate, not a measurement
Anthropic's own API documentation says results 'will not be fully deterministic' even at temperature 0.0 — and its newer Claude models don't accept a temperature setting at all. Google states the same for Vertex AI: a small amount of variation is possible even at temperature 0. Thinking Machines Lab ran one prompt 1,000 times at temperature 0 on one model and got 80 different completions, the most common occurring 78 times. If you ran your brand-tracking prompt once this morning, you sampled one draw from a system its own makers document as non-deterministic — not a fact about the world.
Is your week-on-week change real, or is it noise?
The maths is the NIST formula for the margin of error on a proportion, at the worst case (p = 0.5, the widest possible interval). Find your prompt count, read across.
| Prompts run (n) | 95% margin of error, one measurement | Smallest week-on-week move that isn't just noise | Plain-English verdict |
|---|---|---|---|
| 10 | ±31.0 points | ≈44 points | Don't report this number. Almost any change you see is noise. |
| 20 | ±21.9 points | ≈31 points | Still too noisy for a board deck. Fine for a rough gut-check. |
| 30 | ±17.9 points | ≈25 points | The floor for a monthly number you're willing to defend. |
| 50 | ±13.9 points | ≈20 points | A reasonable target for a recurring client report. |
| 100 | ±9.8 points | ≈14 points | Good enough to catch a real campaign effect, not just a rumour of one. |
| 385 | ±5.0 points | ≈7 points | Survey-grade. Rarely worth the cost for one brand tracker. |
Scroll the table sideways to see every column.
This is the worst-case interval, at p = 0.5. A share that's genuinely near 0% or 100% has a tighter true interval — which is exactly why NIST recommends the Wilson interval over this naive formula for proportions close to the boundary, and why a vendor quoting one flat 'margin of error' for every brand is simplifying more than the maths allows.
What 20 runs actually looks like
Five brands, one prompt, 20 identical runs against the same model. These numbers are constructed to show the effect, not measured from a real account.
The instability is documented, not anecdotal
“CI widths of 5-7 percentage points on citation share are common for SearchGPT domains, and improvements of this magnitude or smaller cannot be reliably attributed to an intervention without repeated sampling and statistical validation.”
<b>Measured citation-instability study, arXiv, 2026</b> — arXiv:2603.08924
Your percentage isn't comparable to theirs, or to last quarter's
Two problems sit underneath the headline number, and neither is about sample size.
The first is counting rules. If a brand is named three times inside one AI answer, does that count once or three times? Vendors don't publish their rule, and the two choices produce different percentages from the identical set of answers. Ask whoever built your tracker which one they use, in writing, before you compare two months of their output.
The second is that engines are not interchangeable denominators. The same underlying study found the median number of sources an engine cites per answer differs by an order of magnitude — roughly 5 to 7 for SearchGPT, 20 to 22.5 for Perplexity, and 36 to 40 for Gemini. A 'share of voice' built from a 6-source answer and one built from a 38-source answer are not the same kind of number, even before the sampling problem. The same paper found domain-level overlap between repeated identical queries as low as 0.29-0.31 for Gemini and only 0.33-0.40 for SearchGPT — so the set of sources feeding your percentage barely repeats itself query to query. Checking which specific sources are actually taking your share tells you more than a blended percentage ever will, and comparing that against your classic Google positions is the fastest way to see whether the gap is an AI problem or a rankings problem wearing a new name.
Four ways this number gets reported wrong
None of these are exotic mistakes. All four are the default behaviour of a spreadsheet, or a cheap tool, left unchecked.
Running 10 prompts once and calling it a baseline
What happens: At n=10 the worst-case margin of error is ±31 points — wider than almost any real change you'll ever see between two periods. A 'baseline' built this way is really just noise with a date stamp on it.
Do this instead: Run at least 30 prompts before you set a baseline, and write the sample size next to the number every time you report it.
Trusting a tool that reports one number with no sample size attached
What happens: A single percentage with no n, no repeat count and no date range is unfalsifiable — you cannot tell if a change is real or if the vendor re-ran a different prompt set.
Do this instead: Ask for the raw counts (named in X of Y runs), not just the percentage. If a vendor won't give you the denominator, they don't have one worth trusting.
Doing it yourself with a spreadsheet and no repeats
What happens: A manual script that runs each prompt once, monthly, has the exact same statistical problem as a paid tool at n=1 — it's free, but it's not more reliable.
Do this instead: The cost trade-off isn't DIY versus paid, it's whether either one runs enough repeats. A script that runs 30 repeats beats a dashboard that runs one, regardless of price.
Comparing your score against a competitor's from a different tool
What happens: Different vendors count multiple mentions differently and sample different prompt sets, so two tools rarely agree on the same brand's number — that disagreement is a methodology gap, not evidence one tool is wrong.
Do this instead: Compare your own number to your own history, from the same tool and the same prompt list. Cross-tool comparison is not a like-for-like measurement.
Five words worth getting right
- Mention
- Your brand named in the visible answer text. No link or source attached required.
- Sample (n)
- The number of prompt runs behind a percentage. A share reported without one is not a measurement.
- Repeat
- Running the identical prompt more than once against the same engine, on the same day, to see how much the answer itself varies before you even change anything.
- Margin of error
- How far a measured percentage could plausibly sit from the true value, at a stated confidence level (usually 95%). Gets wider as n gets smaller, and widest at p = 0.5.
- Wilson interval
- The interval NIST recommends for proportions instead of the naive normal one used in the table above — tighter and more honest near 0% or 100%, which is where most brand-mention shares actually sit.
What to actually do before Monday's deck
Run at least 30 prompts per brand, note the exact count next to the percentage, and don't report a week-on-week move smaller than roughly 1.4 times the margin of error at your sample size. If the board wants a trend rather than a snapshot, a single measurement isn't one — and if you want to know which prompts you should even be testing, your share is only ever measured across the prompts you thought to test. For the fastest sanity check, drill from this aggregate number back into the individual answers it's built from, and check whether your own page is even quotable before you spend the budget chasing the percentage itself.
AI Share of Voice FAQ
What does AI share of voice mean?
What is a good share of voice percentage?
How to get your brand mentioned by ChatGPT?
How to track brand mentions?
How can I check if my brand is mentioned in ChatGPT or other AI tools?
Why did my AI share of voice change overnight?
Can two tools give different share-of-voice numbers for the same brand?
Primary sources used on this page
- Anthropic Messages API: results are not fully deterministic even at temperature 0, and newer models drop the setting entirely — docs.claude.com
- Google Vertex AI: temperature-0 responses are 'mostly' but not fully deterministic — cloud.google.com
- Thinking Machines Lab: 1,000 identical-prompt completions at temperature 0 produced 80 unique outputs — thinkingmachines.ai
- Measured citation instability across engines, including the 5-7 point CI-width finding — arxiv.org
- NIST: the margin-of-error formula for a proportion, and worked sample sizes — www.itl.nist.gov
- NIST: why the Wilson interval is the recommended interval for proportions — www.itl.nist.gov
Check the number underneath the number
A share-of-voice percentage is a starting point, not an answer. These five tell you what's actually behind it.
Read the guide: Is Local SEO Dead? No, but It Has Changed and Do Google Reviews Help SEO? What Google Confirms. Want it handled for you? See our AI SEO (GEO) services.