How we run an AI-visibility audit (our GEO methodology)
Short answer: we run 30 real buyer questions through ChatGPT, Claude, Perplexity, Google AI Overviews, and Gemini on a fixed date, score your share of voice against competitors, document the factual errors AI states about you, hand you a ranked fix roadmap, and re-measure monthly. Every step is run by hand by an expert. It is dated and reproducible, with no black-box score.
When a buyer opens ChatGPT and types "which CRM should I use?" or "best booking software for a small studio," the model doesn't hand back ten blue links. It names a few providers. If yours isn't one of them, there's no second-place click. You get nothing. You're out of the conversation.
Measuring that is harder than it sounds, because AI engines are non-deterministic. Ask the same question twice and you can get two different answers. A screenshot proves nothing, and a single automated "GEO score" out of 100 proves even less. What holds up is a method that is dated, reproducible, and run by a human who knows your category well enough to spot when the AI is confidently wrong about you.
Here's exactly how we run an audit. No black box.
Which questions do you test, and how do you baseline them?
We baseline your visibility by writing roughly 30 questions your actual buyers ask and running each through every major engine on a fixed, logged date. We skip vanity queries with your brand name in them and use the real research prompts: "which provider does X," "alternatives to [competitor]," "is [category tool] worth it for [use case]," "cheapest / most compliant / easiest to integrate."
This is where a human expert earns their keep. A generic scanner doesn't know that "project management tool" and "work OS" surface completely different competitors.
We run each question through ChatGPT, Claude, Perplexity, Google AI Overviews, and Gemini and record who gets mentioned, in what order, and in what tone. Every run is dated and logged, because AI answers drift. That date stamp is what makes the audit reproducible: three months from now you can see where you stand and which direction you're moving. Because engines are non-deterministic, we look at patterns across repeated runs rather than one lucky or unlucky roll, so the baseline reflects what a typical buyer actually sees.
How do you measure our share of voice vs. competitors?
We turn those runs into a leaderboard scoring your brand and each competitor on three numbers:
| Metric | What it answers |
|---|---|
| Mention rate | How often you appear at all when the question is relevant |
| Average rank | When you do appear, are you named first or buried fifth? |
| Sentiment | Is the AI recommending you, hedging, or warning people off? |
This is the slide that changes the room. "We rank on Google" and "AI recommends us" are different games, and plenty of brands that dominate the first are invisible in the second. Seeing a competitor cited in 70% of answers while you sit at 15% makes the problem concrete and fundable.
What does AI get wrong about our brand?
AI engines under-mention brands, and they also make things up. Documenting those errors is the step automated tools miss entirely. We record the specific factual errors and hallucinations the models state about you: wrong pricing, a feature you don't have (or one you do that it denies), an outdated funding status, a compliance claim that's flat wrong, a founder or headquarters that isn't yours.
Catching these needs a human who knows the domain. A bot can't tell that "supports 40 chains" should be 120, or that the AI just confused you with a defunct competitor. An expert reading the transcripts can. A single corrected hallucination about your licensing can be worth more than any ranking bump, because in regulated categories a wrong compliance claim is a deal-killer.
What do we actually do to fix it?
The deliverable is a ranked list of changes, ordered by impact and effort. A score you can't act on is trivia. Most of what moves AI answers is content you already control: your own pages, and your profiles on the third-party sites these engines trust and cite.
Our roadmap typically covers:
- The pages to create or rewrite.
- The specific claims to make crawlable and unambiguous.
- The trusted directories and profiles to correct or claim.
- The freshness updates that matter. Recently updated content gets cited noticeably more often than stale pages ****.
Each item says what to change, why it should move the needle, and roughly how hard it is. You could hand it to your team and start Monday.
How do you know the fixes worked?
Once a month we re-run the same dated question set through the same engines and track the deltas: mention rate up, average rank improving, hallucinations resolved. An audit is a snapshot. AI models update constantly and competitors keep publishing. Monthly re-measure is how you tell whether the fixes worked, and how you catch the day a model update quietly drops you or invents a new error. Same questions, same method, every month: that's what makes the trend line trustworthy.
What you receive
- A dated baseline of 30 buyer questions with the full transcripts.
- A share-of-voice leaderboard: you vs. named competitors on mention rate, rank, and sentiment.
- A documented list of factual errors and hallucinations AI states about you.
- A prioritized, effort-ranked fix roadmap.
- A monthly re-measure so you can see progress, not guess at it.
Every step is run by hand by an expert, dated, and reproducible. That's the whole point. We give you an honest, checkable answer to the only question that matters: when your buyers ask AI which provider to choose, does it name you? We find out, and we make sure it does.