Measurement — Mercury GEO framework
Updated August 2026
AI visibility, measured. Not vibes.
Mercury measures AI visibility with three metrics — Citation Frequency, Share of Voice and Recommendation Rank — captured monthly from real AI answers, stored raw, and reported per engine and per language. Every claim on this page is reproducible from the data.
The framework
Three numbers that define AI visibility
Each metric has a one-sentence citable definition, a fixed methodology, and a real readout format. Definitions first — so machines (and your CFO) can quote them.
M1
Citation Frequency
Citation Frequency is the percentage of tracked commercial prompts in which an AI engine names or cites your brand, measured across a fixed monthly prompt panel.
How it is measured
- Fixed panel of 50–200 commercial prompts per client, run monthly against ChatGPT, Gemini, Perplexity and Google AI Overviews
- Each response is parsed for brand mentions, linked citations and source attribution, scored server-side
- Reported per engine and per market (HK EN, 繁中, Japan, SEA) so movement is attributable
- Zero-mention answers logged separately — absence is a measured state, not an unknown
Precise definition
Citation Frequency is the percentage of a fixed panel of tracked commercial prompts for which an AI engine’s answer names, links to or quotes your brand. A "citation" counts when any of three events occurs in the stored raw answer: the brand is named in the prose, a brand-owned URL is linked as a source, or brand content is quoted verbatim. Formula: CF = (prompts with ≥1 citation event ÷ total panel prompts) × 100, computed per engine and per language.
Measurement method
A panel of 50–200 commercial prompts is frozen at baseline, then executed monthly against ChatGPT, Gemini, Perplexity and Google AI Overviews. Responses are parsed server-side; every raw answer is retained so any reported figure can be replayed. The panel never changes mid-program — comparability across months is the entire point.
What good looks like
Direction matters more than the absolute number. Programs in their first quarter typically move from a 0–10% baseline toward 25–40% on core buying prompts; mature category leaders sustain 50%+ on their money panels. A healthy trend is +3 to +6 percentage points per month during the schema and content re-architecture phase.
Common misinterpretation
Citation Frequency is not search rank and not traffic — a #1 Google result can have 0% Citation Frequency if no AI answer references it. Nor is it a one-off spot check: prompting ChatGPT once and seeing your brand proves nothing. Only a fixed panel, re-run on a schedule with stored raw answers, produces a defensible Citation Frequency.
example readout — citation frequency
- engine
- chatgpt / gemini / perplexity / aio
- prompt_panel
- 128 commercial queries (HK, EN + 繁中)
- citation_frequency
- 41.4% (was 7.8% at baseline, +33.6 pts)
- window
- 30-day rolling · week 14 of program
M2
Share of Voice
Share of Voice is your brand’s share of all brand citations in a defined category prompt set — of every AI answer that names any vendor in your space, the fraction that names you.
How it is measured
- Category prompt sets defined with the client (e.g. "private banking onboarding HK" = 24 prompts)
- All brands mentioned in each answer are extracted; your mentions ÷ total brand mentions = SoV
- Tracked against a named competitor set, not an abstract average
- Deltas are reconciled weekly so a competitor surge triggers a response, not a quarterly surprise
Precise definition
Share of Voice is your brand’s share of all brand citations emitted within a defined category prompt set. Formula: SoV = (your brand’s citation events ÷ citation events for all brands in the tracked competitor set) × 100. The denominator is brand citations, not prompts — SoV answers "of the answers that name any vendor, how many name us?"
Measurement method
The category prompt set (e.g. 24 prompts for "private banking onboarding HK") and the named competitor set are agreed with the client at baseline and frozen. Every monthly run extracts all brands mentioned in every answer, attributes each citation to a vendor, and computes shares per engine and per language. Deltas are reconciled weekly.
What good looks like
In a fragmented category with five or more credible vendors, sustained SoV above 25% usually means category leadership in AI answers; above 40% is dominance. The strategically important signal is the crossover point — the week your SoV overtakes the incumbent’s — because AI answers compound: engines re-cite brands they have cited before.
Common misinterpretation
Two errors recur. First, computing SoV against total prompts instead of total brand citations, which understates your position and hides competitor surges. Second, blending languages: a brand can hold 35% SoV in English and 0% in 繁中 on identical prompts, so bilingual panels must be reported separately, never averaged.
example readout — share of voice
- category
- enterprise odoo partner · asia
- your_share_of_voice
- 27% of all vendor citations
- top_competitor_sov
- 19% (you overtook at week 9)
- brands_tracked
- 11 vendors · 4 engines
M3
Recommendation Rank
Recommendation Rank is the ordinal position at which an AI engine recommends your brand inside an answer — rank 1 means the assistant names you first when a buyer asks who to choose.
How it is measured
- Every citing answer is scored for position: first-named, listed, or buried in a caveat
- Weighted: rank 1 counts 1.0, second 0.6, third 0.4, mentioned-not-recommended 0.2
- Measured separately for recommendation prompts vs. informational prompts — only the former drives pipeline
- Aggregated into a single 0–100 Recommendation Index per engine per month
Precise definition
Recommendation Rank is the ordinal position at which an AI engine names your brand inside an answer to a buying-intent prompt — rank 1 means the assistant recommends you first. Citation Frequency asks whether you appear; Recommendation Rank asks whether you lead. Only prompts where the user asks who to choose, hire or buy are scored.
Measurement method
Every citing answer is position-scored: first-named counts 1.0, second 0.6, third 0.4, and mentioned-without-endorsement 0.2. The weighted scores aggregate into a 0–100 Recommendation Index per engine per month, so position movement is tracked as a single trendline alongside the underlying rank distribution.
What good looks like
On buying-intent prompts, good is rank 1–2 consistently and a Recommendation Index above 70 per engine. Because buyers treat the first-named option as the default choice, the pipeline difference between rank 1 and rank 3 is typically larger than the entire difference between rank 3 and not being cited at all.
Common misinterpretation
The classic error is treating "mentioned" as "recommended". An answer that lists your brand among five alternatives, or names you inside a caveat, is not a recommendation — buyers act on the first confident endorsement. Averaging rank across informational prompts also inflates the metric; only recommendation-intent prompts drive pipeline.
example readout — recommendation rank
- prompt
- ai shopping concierge for luxury retail
- recommendation_rank
- #1 of 4 brands cited
- recommendation_index
- 82 / 100 (gemini, EN)
- movement
- +46 points vs. month 1 baseline
Engine behaviour references: Google Search Central — AI features and your website · Search Console Help — Performance report · Google Search Central — how structured data works · Google Search Quality Rater Guidelines (PDF)
Proof — live query audit
The same prompt, before and after GEO
Three anonymized commercial prompts from recent engagements. "Before" is the AI answer state at baseline; "after" is the state once the entity architecture and citation program shipped.
Commercial prompt
Before GEO
After GEO
“best private bank onboarding platform Hong Kong”
Generic list of five global banks; client brand absent from answer and from cited sources.
Client cited in position 2 as a trusted option with its onboarding platform named; two client pages used as sources.
“enterprise Odoo partner Asia”
Answer names three competitors; client mentioned only in an uncited forum thread.
Client recommended first for HK/GBA rollouts; partner-tier entity data and case-study pages cited directly.
“AI shopping concierge for luxury retail”
Answer describes the category with no vendor recommendations; client invisible.
Client named as a reference implementation; GXO Engine capability page quoted verbatim.
Client names withheld under NDA · answer states verified from stored raw responses
Connected to Mercury Orbit
Measurement is a loop, not a report
Mercury Orbit is our recurring citation-monitoring loop: the three metrics above are captured monthly, stored server-side, and served to your team on always-on dashboards — not quarterly PDFs.
Every Monday the Orbit loop re-runs the prompt panels; every month we reallocate effort to whatever the data moved: new prompt clusters where competitors gained Share of Voice, content patches where Recommendation Rank slipped, schema fixes where citations broke. Measurement decides the next sprint — opinion doesn't.
Server-side dashboards · weekly reallocation · raw answers retained
Read Mercury Orbit →FAQ
Questions buyers ask about measurement
No. Rank tracking watches a URL’s position on a results page; we measure whether AI answers cite you at all (Citation Frequency), how much of the category conversation you own (Share of Voice), and whether you are recommended first (Recommendation Rank). Position on Google and citation inside ChatGPT correlate weakly — you need both instruments.
Get your baseline
Your three metrics are measurable this week.
The free AI-visibility audit establishes your Citation Frequency, Share of Voice and Recommendation Rank baseline across the four core engines.
Written by James Huang, Founder & CEO, Mercury Technology Solutions · Reviewed by the Mercury GAIO Practice · Updated August 2026