Workflow-specific products Content, decks, briefs, proposals, legal, and sales each have a clearer buying path.
Review before delivery Draft, edit, collaborate, approve, and export in the same workspace.
Security + procurement path Security policy, support, and Azure Marketplace buying are public.
Gixo Quill · AEO field guide

LLM visibility scoreboard: how to measure AI citations honestly

An LLM visibility scoreboard should separate four layers: provider-reported citations, controlled synthetic probes, AI-referral visits, and business conversions. Combining them into one opaque score hides provenance. Report the query set, time window, cited domains, and limitations so changes can be reproduced.

First-party measurement: Gixo recorded approximately 6,700 Copilot citations during a measured quarter. The count comes from a Bing AI Performance citation export; it is a citation measurement, not a claim that every cited answer was accurate.

What is LLM visibility?

LLM visibility describes how often and how prominently a brand, domain, or source appears in AI-generated answers for a defined question set. It can mean a mention, a linked citation, a cited domain, or a referral visit, so every report must define which event it counts.

Provider exports are the strongest observation when they contain real citation events. Synthetic probes are useful experiments but reflect the chosen prompts, model, location, account state, and run time. Referral analytics capture only visits that carry detectable referrer information.

Gixo's public number on this cluster is deliberately narrow: approximately 6,700 Copilot citations in a measured quarter from a Bing AI Performance export. It is not presented as total cross-model visibility.

Which LLM visibility signals should stay separate?

DimensionWhat it observesMain limitation
Provider citation exportObserved citations recorded by that provider for the export scopeCoverage and definitions depend on the provider
Synthetic prompt probeWhether a controlled model run mentioned or cited a domainOutputs vary and the prompt set is researcher-selected
AI referral analyticsVisits arriving with a detectable AI-product referrerMany answer views never become visits; attribution can be lost
Conversion analyticsBusiness outcomes attributed to an AI referral or assisted journeyLow volume and multi-touch journeys complicate attribution
Traditional search dataIndexation, impressions, clicks, and queriesIt does not directly report selection inside generated answers

How do you build a reproducible citation scoreboard?

1. Freeze the question set

Version the exact questions, locale, audience, and intent labels. Changing the set changes the denominator.

2. Record provenance

Store provider, model or surface, date, account context, export source, and whether the event is observed or synthetic.

3. Count at multiple levels

Report citations, unique cited queries, cited pages, cited domains, and share of authority instead of one unexplained score.

4. Join referrals carefully

Track AI referrers and conversions as separate downstream layers. Do not infer a visit for every citation.

5. Publish limitations

State missing providers, sampling choices, query bias, and product changes. A reproducible caveat is more useful than false precision.

What does Gixo's citation number prove—and not prove?

The measurement proves that Gixo pages were selected as sources in the cited Copilot/partner query export during the measured period. It does not prove universal visibility, answer accuracy, causation, or future performance. Citation totals vary with query demand, index freshness, product behavior, and the measurement window.

Frequently asked questions

How do you measure LLM visibility?

Freeze a question set first, then measure each layer on its own: pull provider citation exports for observed events, re-run the same prompts on a fixed schedule and label those runs synthetic probes, and track AI-referrer visits in analytics separately again. Record provider, model or surface, locale, account state, and date window next to every number. Do not add the layers together - they have different denominators, so the sum means nothing.

What is an LLM visibility score?

It is a summary of how a brand or domain appears across a defined set of AI answers. A useful score discloses its query set, event definition, provider, and time window.

What is a good LLM visibility score?

There is no industry benchmark, and a vendor quoting one is comparing across question sets that are not the same. The only defensible target is relative: your own citation presence on a frozen question set, measured the same way, moving across successive windows. That is why Gixo publishes roughly 6,700 Copilot citations in a measured quarter as one provider's observed count rather than as a score out of a hundred.

Are AI citations the same as AI referrals?

No. A citation is source inclusion in an answer; a referral is a visit. Many users read a cited answer without clicking.

What is share of authority?

For a controlled query set, share of authority is the proportion of observed citation presence attributed to a domain or brand under a disclosed counting method.

Can prompt tests measure real visibility?

They can measure reproducible samples, not universal exposure. Label them synthetic probes and repeat them consistently.

Why not combine every signal into one score?

Citation exports, probes, referrals, and conversions measure different stages with different denominators. An opaque blend can rise even when the signal you care about falls.

LLM visibility tools: what they measure and what they cannot

Most products sold as LLM visibility tools sample exactly one of the four layers above and present the result as if it covered all four. Before you trust a number, work out which layer the tool is actually watching.

  • Prompt-tracking platforms run a fixed prompt set on a schedule against ChatGPT, Perplexity, Gemini or Copilot and report a mention or citation rate. That is a synthetic probe. It measures the prompts you supplied, from the accounts and locations the vendor runs from, at the moment it ran. It cannot tell you what real users asked.
  • Brand-mention monitors count occurrences of a name in generated text. A mention with no link is not a citation and produces no referral, so mention share and citation share are different denominators and should never be charted together.
  • Provider exports — today that mainly means the AI Performance report in Bing Webmaster Tools — are the only observed citation events on this list. Their limitation is coverage: they describe one provider's surfaces and say nothing about the rest.
  • Referral analytics and server logs count visits that arrive carrying an AI referrer. They systematically undercount, because an answer read without a click leaves no trace at all and some surfaces strip the referrer on the way out.
  • Nothing on the market reports impression share inside AI answers the way Search Console reports it for search results. Any dashboard showing "AI impressions" is modelling that figure from a probe sample, not observing it.

None of this makes the tools useless; it makes their outputs layer-specific. Whichever ones you buy, keep each layer in its own column with provenance attached. A scoreboard is a reporting discipline, not a purchase, and the discipline is what lets you explain a movement three quarters later.

Inside Quill

Start with the amount of structure you already have

Current Gixo Quill quick-create screen with a topic field, visual-plan option, full-writer option, and link to all writing formats
Current Quill quick-create screen. Begin with a topic, open a visual plan, move into the full writer, or browse the complete format library.