Two AI-visibility tools can hand the same brand two different scores. Usually they asked different questions, of different models, a different number of times. This page is our answer: the questions, the engines, the formulas, the error bars — and the list of things we do not claim to measure.
Every brand gets 8 automatically generated questions. A language model writes them from your brand name, your keywords, and one research call that asks an AI engine what your company actually sells — so the questions match your real offer, not only the words you typed into the form.
The generator has three hard rules:
The third rule is the one that decides whether a score means anything. A question that contains your brand name almost always returns your brand. Include those and the number inflates. Every question is tagged branded or non-branded, and the two are reported as separate segments.
Each question also carries a funnel stage (exploring the category / comparing options / ready to pick a provider), a buyer persona and a topic — so a weak score can be traced to a stage instead of staying a single total.
Two extra perception questions per brand ("What is X and what does it do?", "What are the opinions about X — is it a trustworthy company?") are branded by definition and are excluded from every metric. They feed the reputation section only.
You can add your own prompts — up to 50 on Pro and 120 on Agency, counted across the account. Question sets are versioned: the first one activates automatically, and every later change goes through a draft you approve. Your set never changes underneath you.
Which engines we query, and whether the answer may draw on a live web search:
| Engine | Live web search | Reports sources |
|---|---|---|
| ChatGPT (OpenAI) | ✓ | ✓ |
| Gemini (Google) | ✓ | ✓ |
| Perplexity | ✓ | ✓ |
| Claude (Anthropic) | ✗ | ✗ |
| Grok (xAI) | ✗ | ✗ |
Three of the five engines answer with live web search switched on and report the sources they drew on. Claude and Grok answer from their training data. An answer written from training data and one written from live search behave differently, and their numbers should be read knowing which is which.
Starter queries ChatGPT, Gemini and Claude. Pro and Agency query all five.
Model changes are on the record. Each engine is pinned to one model. When a provider retires it, or we repoint an engine, the change is written to your account timeline. A score that moves on the day an engine changed model is labelled as such, instead of looking like a change in your market.
| Engines | Cadence | Repeats per question × engine | Measurements per week | |
|---|---|---|---|---|
| Free check | 3 | once, on request | 1 | 1 |
| Starter | 3 | weekly | 3 on run day | 3 |
| Pro / Agency | 5 | daily | 1 per day | 7 |
We never report a subscription number from a single answer. Weekly plans repeat every question three times on run day; daily plans get their repetitions from consecutive days.
One line. No weighting, no curve, no proprietary index. The denominator is every non-perception answer in the segment. An answer counts once no matter how many times you are named inside it — that is a separate metric, mention frequency.
Counted over mentions rather than answers, so being named twice in one answer counts twice. Only the competitors you track are in the denominator — Share of Voice is relative to the set you named, not to your whole market. Add a competitor and the number moves; that is arithmetic, not a market shift.
Every metric is computed for the whole set, for branded and non-branded questions separately, for each engine on its own, and for each funnel stage. So "43" is never just one number — you can see whether it is one weak engine or one weak stage of the buying journey.
A language model reads every answer and returns, for your brand and for each competitor you track:
For the engines that return sources, we also record every cited URL, its domain, and whether that domain is yours or somebody else's. That is what the source analysis is built from.
LLM answers are not deterministic. Ask the same question twice and the list can come back different. A tool that reports one number from one call is reporting noise.
So we repeat (section 3), then pool the trailing seven days across every completed run and publish a Wilson confidence interval on that rolling window instead of a bare point. A three-point move inside the interval is not news. A move outside it is.
We do not smooth, backfill, or carry a previous value forward to make a chart look calmer.
Blank and zero mean different things, and we keep them apart:
Blank means we do not know. Zero means we measured, and you were not there.
The free check is one run, three engines, one answer per question, no repeats. It is an indication, not the tracked metric.
A subscription queries more engines, more times, and reports a rolling window with a confidence interval. So your free number and your first subscription number can differ — that is the method working, not a bug. Compare subscription windows with subscription windows.
The free report gives you the score, the competitors and the recommendations. The action plan, the source analysis, and mention frequency and context are what a subscription adds.
Every report carries the timestamp of the run it was built from, and every point on a trend chart is a dated run. The cadence for your plan is in section 3.