← All articles

How we run an AI-visibility audit (our GEO methodology)

Short answer: we run 30 real buyer questions through ChatGPT, Claude, Perplexity, Google AI Overviews, and Gemini on a fixed date, score your share of voice against competitors, document the factual errors AI states about you, hand you a ranked fix roadmap, and re-measure monthly. Every step is run by hand by an expert. It is dated and reproducible, with no black-box score.

What is an AI-visibility (GEO) audit? A generative engine optimization (GEO) audit measures whether AI engines mention and recommend your brand when buyers ask "which provider should I use?" It also gives you the specific changes needed to get named more often.

When a buyer opens ChatGPT and types "which CRM should I use?" or "best booking software for a small studio," the model doesn't hand back ten blue links. It names a few providers. If yours isn't one of them, there's no second-place click. You get nothing. You're out of the conversation.

Measuring that is harder than it sounds, because AI engines are non-deterministic. Ask the same question twice and you can get two different answers. A screenshot proves nothing, and a single automated "GEO score" out of 100 proves even less. What holds up is a method that is dated, reproducible, and run by a human who knows your category well enough to spot when the AI is confidently wrong about you.

Here's exactly how we run an audit. No black box.

Which questions do you test, and how do you baseline them?

We baseline your visibility by writing roughly 30 questions your actual buyers ask and running each through every major engine on a fixed, logged date. We skip vanity queries with your brand name in them and use the real research prompts: "which provider does X," "alternatives to [competitor]," "is [category tool] worth it for [use case]," "cheapest / most compliant / easiest to integrate."

This is where a human expert earns their keep. A generic scanner doesn't know that "project management tool" and "work OS" surface completely different competitors.

We run each question through ChatGPT, Claude, Perplexity, Google AI Overviews, and Gemini and record who gets mentioned, in what order, and in what tone. Every run is dated and logged, because AI answers drift. That date stamp is what makes the audit reproducible: three months from now you can see where you stand and which direction you're moving. Because engines are non-deterministic, we look at patterns across repeated runs rather than one lucky or unlucky roll, so the baseline reflects what a typical buyer actually sees.

How do you measure our share of voice vs. competitors?

We turn those runs into a leaderboard scoring your brand and each competitor on three numbers:

MetricWhat it answers
Mention rateHow often you appear at all when the question is relevant
Average rankWhen you do appear, are you named first or buried fifth?
SentimentIs the AI recommending you, hedging, or warning people off?

This is the slide that changes the room. "We rank on Google" and "AI recommends us" are different games, and plenty of brands that dominate the first are invisible in the second. Seeing a competitor cited in 70% of answers while you sit at 15% makes the problem concrete and fundable.

What does AI get wrong about our brand?

AI engines under-mention brands, and they also make things up. Documenting those errors is the step automated tools miss entirely. We record the specific factual errors and hallucinations the models state about you: wrong pricing, a feature you don't have (or one you do that it denies), an outdated funding status, a compliance claim that's flat wrong, a founder or headquarters that isn't yours.

Catching these needs a human who knows the domain. A bot can't tell that "supports 40 chains" should be 120, or that the AI just confused you with a defunct competitor. An expert reading the transcripts can. A single corrected hallucination about your licensing can be worth more than any ranking bump, because in regulated categories a wrong compliance claim is a deal-killer.

What do we actually do to fix it?

The deliverable is a ranked list of changes, ordered by impact and effort. A score you can't act on is trivia. Most of what moves AI answers is content you already control: your own pages, and your profiles on the third-party sites these engines trust and cite.

Our roadmap typically covers:

  • The pages to create or rewrite.
  • The specific claims to make crawlable and unambiguous.
  • The trusted directories and profiles to correct or claim.
  • The freshness updates that matter. Recently updated content gets cited noticeably more often than stale pages ****.

Each item says what to change, why it should move the needle, and roughly how hard it is. You could hand it to your team and start Monday.

How do you know the fixes worked?

Once a month we re-run the same dated question set through the same engines and track the deltas: mention rate up, average rank improving, hallucinations resolved. An audit is a snapshot. AI models update constantly and competitors keep publishing. Monthly re-measure is how you tell whether the fixes worked, and how you catch the day a model update quietly drops you or invents a new error. Same questions, same method, every month: that's what makes the trend line trustworthy.

What you receive

  • A dated baseline of 30 buyer questions with the full transcripts.
  • A share-of-voice leaderboard: you vs. named competitors on mention rate, rank, and sentiment.
  • A documented list of factual errors and hallucinations AI states about you.
  • A prioritized, effort-ranked fix roadmap.
  • A monthly re-measure so you can see progress, not guess at it.

Every step is run by hand by an expert, dated, and reproducible. That's the whole point. We give you an honest, checkable answer to the only question that matters: when your buyers ask AI which provider to choose, does it name you? We find out, and we make sure it does.

Frequently asked questions

Can I use an automated GEO scanner instead?
Automated scanners give you a number. They can't tell when an AI states something factually wrong about your brand, and they don't know your category well enough to write the questions your buyers actually ask. Our audit is run by hand by an expert who catches hallucinations and gives you an actionable roadmap you can start on Monday.
AI answers change every time you ask. How can an audit be reliable?
That's why we use a dated, reproducible method rather than a one-off screenshot. We run a fixed set of 30 questions across engines, look at patterns over repeated runs rather than single answers, and stamp everything with a date. That way the baseline reflects what a typical buyer sees, and the monthly re-measure shows a real trend.
Which AI engines do you check?
ChatGPT, Claude, Perplexity, Google AI Overviews, and Gemini. These are the engines B2B buyers are actually using to research providers. We record who gets mentioned, in what order, and in what tone on each.
How many questions do you test, and who writes them?
30 real buyer questions, written by us based on how people in your category actually research. That includes alternatives-to-competitor and use-case queries, going well beyond prompts with your name in them. We learn your category first, so the questions surface the competitors your buyers actually compare you against.
What do I actually get at the end?
A dated baseline with full transcripts, a share-of-voice leaderboard against named competitors, a documented list of factual errors AI makes about you, a prioritized fix roadmap ranked by impact and effort, and a monthly re-measure to track progress.
Is this a one-time report or ongoing?
Both are available. The initial audit is a full baseline and roadmap. Because AI models update constantly and competitors keep publishing, most clients keep the monthly re-measure so they can confirm fixes are working and catch new errors or ranking drops early.
Does this work for any kind of business?
Yes. The method works for any B2B or B2C brand whose buyers research on AI. A human expert learns your category and its vocabulary first, so the same five-step approach applies across every market we run it in.

Want to know if AI recommends you?

Get an expert-run AI-visibility audit. See whether ChatGPT, Claude, Perplexity and Google AI Overviews name your brand, and how to fix it.