How do you measure AI visibility?
The benchmark method: fixed search queries, a baseline measurement across multiple platforms and a monthly re-measurement where the delta is the evidence.
Lees dit in het Nederlands →$ measure --method=benchmark --interval=monthly
How do you measure AI visibility?
You measure AI visibility by fixing a set of search queries in advance, noting for each query whether your business appears as a source in ChatGPT, Perplexity, Claude and Google AI Overviews, and repeating that measurement monthly with the same queries. The difference from the first measurement is your evidence. This piece describes the method; you can carry it out entirely yourself, without us.
A lot is claimed about visibility in AI answers and little is measured. That's no coincidence: answers differ per session, per phrasing and per platform, and in that noise everyone can find proof they're right. The solution isn't measuring more, but first agreeing what you measure. The method in four steps:
- Fix a set list of search queries that your target client actually asks, before you measure.
- Agree on what counts as visible: mentioned as a source, linked, or recommended as a party.
- Run the baseline measurement: every query through every platform, with the date and method noted.
- Re-measure monthly with the same queries; the difference from the baseline measurement is your evidence.
The sections below elaborate on each step.
01 / the benchmarkFix the queries before you measure
The benchmark is a fixed list of search queries that your target client actually asks, set before the first measurement runs. Twenty to thirty queries is workable: enough to see a pattern, few enough to sustain monthly.
Also fix what counts as visible. Are you mentioned as a source with a link, mentioned without a link, or recommended as a party. Those are three different outcomes, and anyone who mixes them up is comparing apples with pears next month. You get the queries themselves from conversations with clients, from your inbox and from the search terms your site is already shown for.
02 / the baseline measurementRecording the first state
The baseline measurement is the first full round: every query through every platform, noting what happens for each combination. Do you appear, who does appear, and with which page. Note the date and method alongside it, because only then can a later measurement be fairly compared.
Note three things per answer: whether your business appears, which parties do get mentioned, and which page of those parties is cited. That third point is the most instructive, because it shows what kind of page carries the answer: a knowledge article, a comparison, a review page. That's where your improvement list lies.
Expect a confronting outcome. Anyone doing this for the first time often turns out to be visible on almost no query, while a handful of competitors and comparison sites keep coming back. That's not bad news but a starting point: now you know for certain, instead of suspecting it. How to do such a first check yourself in ten minutes is in how to check whether your business appears in ChatGPT.
03 / the re-measurementThe delta is the evidence
Every month the same measurement runs again: the same queries, the same platforms, the same definition of visible. The only figure that matters after that is the delta: on how many queries are you now visible where that wasn't the case before.
Without a benchmark set in advance, any result can be defended after the fact, and is therefore worth nothing.
The delta forces honesty in two directions. If something moves, it's there in black and white and it's not a sales pitch. If nothing moves, that too is visible, and then the right question is why, instead of it staying hidden behind loose success stories.
If you work with an agency, this is also your control instrument. Ask for the benchmark and the baseline measurement before the start, and for the re-measurement in every monthly report after that. A party willing to do this has nothing to hide; a party that avoids it is asking you to pay for something you'll never be able to check.
04 / the meansTools or manual work
Tools exist that automate this work, such as Otterly and Peec AI: they run your queries periodically through multiple platforms and keep track of the score. For anyone tracking many queries or multiple brands, that's the logical route. Outsourcing is also an option; what such an outsourced baseline measurement looks like is on the page about the scan.
Manual work works too, and for a first stretch it's often enough: a spreadsheet with the queries in the rows, the platforms in the columns, and one fixed moment per month when you run the round. For 25 queries across four platforms, reckon on half a day per month.
| Tools such as Otterly and Peec AI | Manual work with a spreadsheet | |
|---|---|---|
| How it works | Run your queries periodically through multiple platforms and keep track of the score | Queries in the rows, platforms in the columns, one fixed measurement moment per month |
| What it requires | Setting up and maintaining the query list | Half a day per month for 25 queries across four platforms |
| Best for | Tracking many queries, or multiple brands at once | A first stretch and one brand |
The recommendation for those starting out: do the manual work for the first months. The method determines the value of the measurement, not the tool, and anyone tired of the manual work after three months will by then know exactly what a tool needs to do for them.
05 / the limitWhat this measurement can't do
A snapshot remains a snapshot. AI answers vary per session, so one isolated measurement proves little; the pattern over months is what counts. So measure fixed queries at a fixed moment, and don't draw conclusions from one deviating round.
And the measurement says nothing about revenue. Being visible in an answer is not the same as getting a call. Anyone who wants to know what visibility yields should also track a traffic and inquiry figure alongside this measurement. For businesses without clients searching, this whole instrument is superfluous, and that may also be the conclusion.
What is a benchmark in AI visibility?+
How many search queries do you need for a baseline measurement?+
Why do you need to measure again every month?+
Do you need a tool to measure AI visibility?+
What does the measurement say about my revenue?+
Would you rather have the baseline measurement outsourced, with a report and improvement points included?
The scan
This is the baseline measurement: fixed search queries on ChatGPT, Gemini, Claude, Perplexity and Google. Current prices are on the pricing page. Even without a follow-up, you keep the report.
This text was produced with AI assistance and checked and approved by a human before publication.