AI search visibility tracking means measuring how often, and in what terms, AI assistants mention your brand when people ask them buying questions. It is not rank tracking with a new label. There are no positions to record, most surfaces publish no analytics at all, and the one surface with the largest reach — Google AI Overviews — deliberately does not break out its own appearances in Search Console. Any honest measurement programme has to start by naming what is knowable and what is not.
This guide covers exactly that: what each major surface exposes, what it does not, why prompt-based citation testing is the only method that works across all of them, and what a sane cadence looks like so you are measuring a trend instead of chasing noise.
What "visibility" means when there is no ranked list
On a classic results page, visibility is ordinal. You are somewhere between one and one hundred, and moving up is the job.
In an assistant answer, visibility is compositional. The model produces one response assembled from a small set of retrieved sources plus its own parameters. You are either part of that response or absent from it. There is no consolation prize for being the eleventh most relevant brand.
So the measurable quantities change:
- Presence rate. Across a fixed panel of prompts, in what share of answers is your brand named at all?
- Citation rate. In what share is one of your URLs actually linked as a source? Presence without citation is common and worth separating.
- Share of voice. When you are named, who else is named in the same answer, and how often relative to you?
- Breadth. How many distinct pages of yours get cited, not just how many citations you accumulate.
- Sentiment and accuracy. Is the description of your product correct? An assistant confidently repeating an outdated price or a feature you removed is a visibility problem with a support-ticket tail.
That last set is what turns raw counts into AI strategic visibility: knowing not only whether you are present, but whether the version of you being presented to buyers is the one you would write yourself.
Surface by surface: what you can and cannot see
Microsoft Copilot and Bing
The best-instrumented surface, by a wide margin. Bing Webmaster Tools has a Performance section with an AI Performance report showing citation counts over a date range, plus a Pages tab listing which of your URLs were cited.
Copilot also has no separate crawler — it rides bingbot — so your Bing index coverage and your robots.txt rules for Bingbot directly determine your eligibility. That makes the whole chain inspectable end to end, which is unusual. We cover the tracking workflow and one nasty robots.txt trap in Copilot rank tracking.
What you still cannot see: the prompts that produced the citations. You get counts and URLs, not questions.
Google AI Overviews
Be blunt about this one, because a lot of tooling is vague about it: Google Search Console does not break out AI Overview appearances. There is no AI Overview search-appearance filter, no separate impressions line, and no citation report. Clicks and impressions from pages that appear inside an AI Overview are folded into your regular web search totals with no way to isolate them.
So you cannot measure AI Overview presence from Search Console. Anyone telling you otherwise is inferring, not measuring. What you can do:
- Watch for the classic symptom pattern — impressions steady or rising while clicks fall on informational queries — and treat it as a hypothesis, not a finding.
- Sample manually or programmatically: run your target queries and record whether an AI Overview appears and whether you are cited in it.
- Segment by query intent, since AI Overviews trigger far more on informational than transactional queries.
An inference clearly labelled as an inference is useful. An inference presented as a metric will eventually be used to justify a bad decision.
ChatGPT
No publisher analytics. OpenAI does not provide a webmaster console, citation report or impression data. When browsing is used, answers include links, and those visits show up in your own analytics as referral traffic — which is real data, but it measures clicks, not mentions. Most assistant answers never produce a click at all, so referral traffic systematically under-counts your actual presence.
The workable method here is direct: ask it, on a schedule, with a stable prompt panel, and record the answers.
Perplexity
Answers are citation-dense and the sources are visible in the response, which makes manual and programmatic observation straightforward. There is still no publisher-side console, so measurement is again observational rather than reported. Referral traffic appears in analytics and is worth segmenting, with the same caveat that clicks understate mentions.
The pattern
One surface reports to you (Bing). One surface deliberately does not separate its AI layer (Google). The rest report nothing. That asymmetry is the reason prompt-based testing exists — it is the only method that produces comparable numbers across all of them.
Prompt-based citation testing, in practice
The method is simple and the discipline is where it lives or dies.
1. Build a prompt panel that reflects buying behaviour. Twenty to forty prompts, written the way a prospect types, not the way a keyword tool exports. Cover five shapes:
- Category discovery: "what tools do X"
- Comparison: "A vs B for a small team"
- Recommendation: "best X for Y situation"
- Objection: "is X worth it", "cheaper alternatives to X"
- Branded: "what is [your brand]", "is [your brand] any good"
The branded set matters more than people expect. It is where you find out whether assistants describe you accurately.
2. Freeze the panel. Changing prompts changes the measurement. Add new prompts as a separate cohort; never edit the existing ones mid-quarter.
3. Run every prompt on every surface you care about, on a schedule. Same prompts, same order, logged.
4. Record structured fields, not screenshots. For each run: brand mentioned (yes/no), own URL cited (yes/no, which URL), competitors named (list), any factual claim about your product (verbatim). Structured fields are what let you compute a trend three months later.
5. Repeat before concluding. Assistant answers vary run to run. A single miss is not a regression. Compare weekly aggregates, not individual answers.
6. Control the neutral state. Run without any personalisation or account context where possible, so you are measuring the default answer rather than one shaped by your own history.
This is what AutoSEO's AI-visibility tracking automates: we query language models with buyer prompts on a schedule and log whether your brand is cited, together with competitor share-of-voice in the same answers. Surface coverage keeps changing as these platforms change, so check the current list rather than assuming a specific assistant is included — see pricing, or run a one-off check with the free AI visibility checker.