Share of voice has always been a stand-in for attention. In advertising it is your spend against the category's. In PR it is your mentions against everyone else's. The logic has held for decades: the more of the conversation you hold, the more likely a buyer thinks of you first. AI answers are a new place to run that same measurement, and most teams have not started.
What the number actually is
AI share of voice is the share of brand mentions in a set of AI answers that belong to you. Take the questions your buyers ask, run them across the models your buyers use, count every brand each answer names, and your slice of that total is your share of voice.
Say forty questions produce forty answers, and those answers name brands two hundred and ten times between them. Thirty of those mentions are yours, so your share of voice is fourteen percent. Your nearest competitor holds fifty-two, a shade under a quarter. That arithmetic asks only that someone write down every name in every answer, and that is the part teams skip.
It travels with a second number worth keeping next to it. Visibility is the share of answers that mention you at all, regardless of who else appeared. A brand can hold a high visibility and a low share of voice: it gets named in most answers, always last, in a list of twelve. Those two numbers pull apart in useful ways, so record both.
Average position is the third. An answer names brands in an order, and being third in a list of four reads very differently from being first. Most of the time the three move together. When they diverge, the reason is usually interesting.
Why it resists being gamed
Marketers meet a new metric and immediately ask how to juice it. This one is stubborn, and the reasons it is stubborn are the reasons it is worth tracking.
- It draws on many sources at once, so a single press release barely registers against everything else the model has read.
- It tracks the description that has settled around you, so substance counts for more than budget.
- It moves slowly, so gains you earn tend to stay.
That gives you something that behaves like brand equity. It is hard to spike on purpose, which is exactly why a steady climb means something.
The three things that will fool you
A share-of-voice number is only as good as the run behind it. Three conditions turn a real measurement into a misleading one, and all three are invisible unless you record them alongside the number.
First, which models answered. A rate measured on cheap auto-routed models is a different quantity from one measured on the assistants your customers actually open. Both are legitimate; averaging them together and calling the result your share of voice is not. Write down the model behind every data point.
Second, whether the answer retrieved anything. Some models search the web on every call, some only when asked, and some cannot search at all. A model that did not retrieve is answering from training data, so the URLs it names came out of its memory, and a page you published last month cannot appear. Source reports built on ungrounded answers describe a memory, and acting on them means chasing citations that were never made.
Third, how much of the run landed. If you scheduled six hundred questions and two hundred and fifty produced an answer, the rest failed on rate limits, timeouts or provider outages. Your share of voice then sits on a sample that skews toward whichever model stayed up. A measurement window with low coverage looks like a visibility drop and is an outage. Check coverage before you go looking for a cause in your marketing.
Any one of the three moves a number further than your marketing did. So record the model, the retrieval status and the coverage next to every figure, and read a change only against a window that matches on all three. A drop from nineteen percent to eleven means something when both windows ran on the same four assistants with the web switched on. It means nothing when the second window lost half its questions to rate limits. When you cannot tell those apart, rerun the window before you explain it.
How to start keeping score
You do not need a research budget. Pick twenty to fifty questions that match how your buyers describe their problem. Run them across ChatGPT, Claude, Gemini and Grok on a regular schedule. Record every brand each answer names, the order, the model, and whether the answer retrieved anything. Your share of those mentions, watched over time, is the number. A single snapshot tells you almost nothing; the trend line is the whole point.
Do this by hand and it becomes a spreadsheet that is stale within a week and abandoned within a month. Automate it and it becomes a leading indicator you can steer by. That is what Brandtrace runs for you: daily tracking across the major models, broken out by competitor, with the model and the retrieval status attached to every number so you can tell a real drop from a bad run. Related reading: five ways to improve how models describe you.