GEO Audit Walkthrough: The Complete Mercury Method (2026, With Real Scores)
Most GEO audits are delivered as a mystery. A deck lands, a score appears, a list of fixes follows — and the buyer has no way to know whether the AI visibility score was measured or invented. The agency has every incentive to keep the method opaque: if you can't reproduce the audit, you can't dispute it, and you can't leave.
We run the opposite play. This is our complete GEO audit methodology — every step, every metric, how the score is computed, and what the numbers looked like on five real audits. 百聞不如一見 — one look beats a hundred descriptions. If a prospective client reads this and runs the audit themselves, good. The methodology being reproducible is the product.
TL;DR: A GEO audit answers two questions most reports blur into one. Method A: when AI engines verify claims about you, can they corroborate them? Method B: when buyers ask open questions, does AI select you unprompted? We score both 0–100, fuse them into a unified score, and report the gap — because a brand can score well on corroboration and still lose every open query, and that gap is the actual diagnosis. Below: the seven steps, the scoring, and real anonymized baselines from a Hong Kong lender (62), an insurer (gap closed 24→4 in 30 days), a luxury resort (70), and a cross-border property agency (50) — plus Mercury's own trajectory from 69 to 82.
I am James, CEO of Mercury Technology Solutions. We build AI visibility systems for enterprises — audits, entity infrastructure, always-on monitoring — the measurable half of enterprise digital transformation. Everything in this walkthrough is the same process we run for paying clients, minus the client names.
What a GEO Audit Actually Measures
A GEO audit is not an SEO audit with new acronyms. SEO asks whether Google can rank your pages. GEO asks whether ChatGPT, Perplexity, Gemini and Claude can understand, trust, and cite your brand when they assemble an AI answer.
Two different capabilities, two different scores:
Unified Score = Corroboration (Method A) × Selection (Method B) — and the Gap between them is the diagnosis.
Method A — Corroboration. Can AI verify what you claim? The mechanical layer: crawler access, structured data, entity consistency, extractable content. Borrowable, buildable, fast.
Method B — Selection. Does AI choose you when nothing forces it to? Open queries — "who's the best X for Y?" — are answered from the model's priors about entities, not from placements. You cannot rent this score. You build it or you don't have it.
The gap tells you which problem you have. High A, low B: visible when named, absent when it matters — a rented strategy. Low A, decent B: the market already talks about you, but the machines can't verify what they hear — wasted authority. Equal and low: early days, build the floor.
The Walkthrough: Seven Steps
Step 1. Build the query matrix. Ten to thirty buyer prompts, split into two classes. Forced queries name the brand ("[brand] vs [competitor]", "is [brand] licensed") — these are defenses. Open queries name only the need ("best second-mortgage lender in Hong Kong", "which insurer covers pre-existing conditions") — these are where buyers actually are. The matrix is scoped with the client: category, sub-categories, competitor set of three to five, both languages where the market is bilingual. The matrix is the audit's contract — everything after this step runs against it.
Step 2. Run the baselines. Every prompt goes through the live AI engines — ChatGPT, Perplexity, Gemini, Claude. For each answer we record: brand mention (yes/no), citation and its source URL, position in the answer, which competitors appeared, and any factual errors or hallucinations about the brand. This produces the "before" snapshot — the most persuasive page in the deliverable, because executives trust what they can scroll.
Step 3. Method A — the corroboration audit. Technical reality, checked by hand: Are GPTBot, ClaudeBot, PerplexityBot and Google-Extended explicitly allowed in robots.txt — or silently blocked? Is the site server-rendered, or is the catalog trapped in a JavaScript shell no crawler executes? We audited one site that ships a genuine per-article Markdown pipeline and an llms.txt — advanced, rare — wrapped inside a 3KB client-side-rendered shell that GPTBot cannot read. Advanced signals on an unreadable page score zero. Then: schema graph completeness, entity description convergence, content extractability (does the direct answer sit in the first hundred words, or under a warm-up paragraph?), freshness signals, sitemap integrity.
Step 4. Method B — the selection audit. The entity layer: a consistent, machine-agreeable description of the company — Wikidata present, Wikipedia current, canonical blurb locked, the same facts in the same words across site, schema, and independent surfaces. Reviews and community: does anyone unaffiliated say anything about you, anywhere the AI models read? And the voice map — the newest module: which named humans do the engines cite when explaining this category, and is the client's expertise among them? AI cites people, not only pages; a category whose answers quote five analysts is a category with five doors, and we map every one.
Step 5. Benchmark the competitors. The same matrix, the same engines, run against each competitor. This converts absolute scores into position: a 62 means one thing where the leader sits at 58, another against a leader at 84. Competitor gaps also expose white space — sub-categories where nobody is established in AI answers yet, the cheapest ground to take.
Step 6. Score and diagnose. Method A and Method B each land 0–100 across sixteen dimensions (five live queries per engine anchor the selection half; technical and entity checks fill the corroboration half). The unified score fuses them; the gap — A minus B — is the headline number. We report score, gap, rank among the benchmarked set, and the three failure patterns driving the largest point losses.
Step 7. The fix plan and the 90-day re-audit. Findings become a P0/P1 sequence: P0 items are same-week mechanical fixes (crawler whitelist, schema repair, description lock); P1 items are 30/60/90 structural work — entity building, review surfaces, voice-map capture, content restructure. Then the discipline that separates GEO from one-page SEO reports: AI models change quarterly, so the audit is quarterly. Every engagement ends with a re-audit date, not a farewell.
The Receipts: Five Real Audits
What the AI visibility numbers look like in practice — all real, all anonymized:
A licensed Hong Kong lender: 65 / 60 / unified 62. AI already cites them on three of five forced buyer queries — citation presence without entity infrastructure. Stale 2021 Wikipedia, no Wikidata, no consistent description. The machines quote them and can't say who they are. Diagnosis: borrowed authority, unbuilt entity layer.
A Hong Kong insurer: unified 58, gap 4. A 30-day follow-up: the gap had been 24. The fix sequence closed it six-fold in a month — but splitting by product line revealed Travel scoring 64 while VHIS sat at 42. Brand averages lie; product-level audits don't.
A luxury resort: 69 / 70 / unified 70. Niche champion — cited third on "best luxury family hotel" queries behind two global flagships. The ceiling isn't technical; it's award-gated. Some gaps are bought with years, not sprints — the audit says which.
A cross-border property agency: 59 / 44 / unified 50, gap +15. The signature selection failure: genuinely advanced content signals, invisible on every open AI query. Right machinery, wrong answers.
Mercury ourselves: 69 → 80 → 82 across three audits. We run our own method on our own brand first. Trust an agency that publishes its own score trend.
What a GEO Audit Costs
The market prices standalone GEO audits between roughly $1,500 and $7,500, depending on query-matrix size, engine coverage, and language count. Ours is scoped the same way — matrix size and market complexity — with one structural difference: we treat the baseline audit as the first cycle of a quarterly cadence, because a score that isn't re-measured is a photograph of a moving object. The audit either becomes a monitoring rhythm or it was entertainment.
Why We Publish the Method
Rommel's rule: time spent on reconnaissance is never wasted. The audit is reconnaissance — and a methodology you can inspect is one you can act on. Every framework in this post — the query matrix, Method A/B, the gap, the voice map — is named so it can be cited, tested, and argued with. That is not generosity. A method the machines can extract is a method that gets extracted — and an agency that shows its work becomes part of the AI answer.
Stop buying scores you can't reproduce. Start demanding the method with the number.
Mercury Technology Solutions: Accelerate Digitality.
Originally published on MTS Blog & Research