1. Freeze the question set
Version the exact questions, locale, audience, and intent labels. Changing the set changes the denominator.
An LLM visibility scoreboard should separate four layers: provider-reported citations, controlled synthetic probes, AI-referral visits, and business conversions. Combining them into one opaque score hides provenance. Report the query set, time window, cited domains, and limitations so changes can be reproduced.
LLM visibility describes how often and how prominently a brand, domain, or source appears in AI-generated answers for a defined question set. It can mean a mention, a linked citation, a cited domain, or a referral visit, so every report must define which event it counts.
Provider exports are the strongest observation when they contain real citation events. Synthetic probes are useful experiments but reflect the chosen prompts, model, location, account state, and run time. Referral analytics capture only visits that carry detectable referrer information.
Gixo's public number on this cluster is deliberately narrow: approximately 6,700 Copilot citations in a measured quarter from a Bing AI Performance export. It is not presented as total cross-model visibility.
| Dimension | What it observes | Main limitation |
|---|---|---|
| Provider citation export | Observed citations recorded by that provider for the export scope | Coverage and definitions depend on the provider |
| Synthetic prompt probe | Whether a controlled model run mentioned or cited a domain | Outputs vary and the prompt set is researcher-selected |
| AI referral analytics | Visits arriving with a detectable AI-product referrer | Many answer views never become visits; attribution can be lost |
| Conversion analytics | Business outcomes attributed to an AI referral or assisted journey | Low volume and multi-touch journeys complicate attribution |
| Traditional search data | Indexation, impressions, clicks, and queries | It does not directly report selection inside generated answers |
Version the exact questions, locale, audience, and intent labels. Changing the set changes the denominator.
Store provider, model or surface, date, account context, export source, and whether the event is observed or synthetic.
Report citations, unique cited queries, cited pages, cited domains, and share of authority instead of one unexplained score.
Track AI referrers and conversions as separate downstream layers. Do not infer a visit for every citation.
State missing providers, sampling choices, query bias, and product changes. A reproducible caveat is more useful than false precision.
The measurement proves that Gixo pages were selected as sources in the cited Copilot/partner query export during the measured period. It does not prove universal visibility, answer accuracy, causation, or future performance. Citation totals vary with query demand, index freshness, product behavior, and the measurement window.
Freeze a question set first, then measure each layer on its own: pull provider citation exports for observed events, re-run the same prompts on a fixed schedule and label those runs synthetic probes, and track AI-referrer visits in analytics separately again. Record provider, model or surface, locale, account state, and date window next to every number. Do not add the layers together - they have different denominators, so the sum means nothing.
It is a summary of how a brand or domain appears across a defined set of AI answers. A useful score discloses its query set, event definition, provider, and time window.
There is no industry benchmark, and a vendor quoting one is comparing across question sets that are not the same. The only defensible target is relative: your own citation presence on a frozen question set, measured the same way, moving across successive windows. That is why Gixo publishes roughly 6,700 Copilot citations in a measured quarter as one provider's observed count rather than as a score out of a hundred.
No. A citation is source inclusion in an answer; a referral is a visit. Many users read a cited answer without clicking.
For a controlled query set, share of authority is the proportion of observed citation presence attributed to a domain or brand under a disclosed counting method.
They can measure reproducible samples, not universal exposure. Label them synthetic probes and repeat them consistently.
Citation exports, probes, referrals, and conversions measure different stages with different denominators. An opaque blend can rise even when the signal you care about falls.
Most products sold as LLM visibility tools sample exactly one of the four layers above and present the result as if it covered all four. Before you trust a number, work out which layer the tool is actually watching.
None of this makes the tools useless; it makes their outputs layer-specific. Whichever ones you buy, keep each layer in its own column with provenance attached. A scoreboard is a reporting discipline, not a purchase, and the discipline is what lets you explain a movement three quarters later.
Inside Quill