To track your brand’s visibility in AI assistants, write a fixed set of 30 to 50 realistic prompts, run them on the assistants your customers actually use, and log whether you are mentioned, cited, described accurately and placed next to competitors. Repeat on a schedule and judge trends rather than single answers, because outputs change from run to run.
AI answers are now part of ordinary search: Pew Research Center reported in October 2025 that 65% of U.S. adults at least sometimes come across AI summaries in search results. If customers see those answers, you need a way to know what they say about you.
Why this is not the same as rank tracking
A rank tracker answers a simple question: where does this URL sit for this keyword? An AI assistant does not return a list of ten positions. It writes an answer, may or may not name brands, may or may not show sources, and can write a different answer the next time you ask. So the unit you measure changes. Instead of a position, you record presence (were you named), attribution (was your site cited), accuracy (was what it said true) and context (who else was named beside you).
That is why I treat AI visibility tracking as sampling, not as a number. One answer proves almost nothing. A pattern across many prompts, several assistants and several weeks tells you something you can act on. If you want the wider strategy this measurement serves, start with my guide to generative engine optimisation and come back here to set up the scoreboard.
Step 1: Build a prompt set that mirrors real buying questions
Your prompt set is the whole instrument, so build it carefully. Do not write prompts the way a marketer would. Write them the way a customer who does not know your brand would ask. Pull wording from sales calls, support emails, the questions people type into your site search, and the queries that already bring impressions in Search Console.
Group the prompts into four buckets so you can see where you are strong and where you are invisible:
- Category prompts: “What are the best accounting tools for a small import business in Karachi?” You are hoping to be named without being mentioned in the question.
- Comparison prompts: “Compare [your brand] and [competitor] for a clinic with three branches.”
- Problem prompts: “How do I stop my online shop losing customers at checkout?” These show whether your content is used as an explanation.
- Brand prompts: “What is [your brand] and who is it for?” These check accuracy, which is where most surprises appear.
As a starting point I use 30 to 50 prompts in total, weighted towards category and problem prompts, because those are the moments where a new customer meets a brand for the first time. That number is a working rule, not a standard. Fewer than about 20 and a single odd answer will swing your picture; many more and the spreadsheet becomes a chore nobody maintains. Lock the wording once you are happy. If you keep editing prompts, you cannot compare this month with last month.
Step 2: Choose which assistants to sample
Sample where your customers are, not where the tools are fashionable. Ask a few real customers which assistants they use before you decide. For most businesses the sensible starting list is ChatGPT, Gemini, Perplexity, Microsoft Copilot and Google’s own AI features in search. Different assistants behave differently, so record the conditions each time.
| Surface | What to record | Caution |
|---|---|---|
| ChatGPT | Whether web search was used, brands named, links shown | An answer from model memory and an answer using live search can differ completely |
| Gemini | Brands named, sources if shown | Behaviour depends on the mode and account you use |
| Perplexity | Named brands, the sources listed, order of sources | Check the cited pages yourself, do not assume the summary matches them |
| Microsoft Copilot | Brands named, sources if shown | Note which Copilot experience and mode you used |
| Google AI Overviews and AI Mode | Whether an overview appeared, whether you were linked | Not every query triggers one, so “no overview” is a valid result to log |
Google’s own documentation says there are no additional requirements to appear in AI Overviews or AI Mode beyond being indexed and eligible to show with a snippet, and that traffic from these features is included in overall Search Console data (see Google Search Central on AI features). That matters for tracking: Search Console will not give you a separate “AI visibility” report, so manual sampling is how you see the answer text itself. For how each assistant finds and cites sources, see my posts on ChatGPT search SEO, Perplexity SEO and Google AI Mode.
Step 3: Record the same fields every time
A spreadsheet with one row per prompt per assistant per run is enough. These are the columns I would include:
- Date and run number. Run each prompt three times in a session so you can see the spread.
- Prompt ID and bucket. Keeps the wording fixed and lets you filter by bucket.
- Assistant and mode. For example, with or without web search.
- Mentioned (yes or no). Was your brand named in the answer text?
- Cited (yes or no, and which URL). Was a page of yours listed as a source?
- Position in the answer. First named, among several, or last mention. Keep it rough.
- Sentiment. Positive, neutral, negative or mixed. Add the phrase it used in one cell so the label can be checked later.
- Accuracy. Correct, partly wrong or wrong, with a note on what was wrong (old pricing, wrong location, a service you do not offer).
- Competitors named. A simple list.
- Sources cited that are not yours. This column becomes your outreach and PR list.
Save the full answer text or a screenshot for any row you may need to show someone later. Without the evidence, an argument about what the assistant “said” turns into a memory contest.
Why results vary from run to run
If you ask the same prompt twice and get two different answers, nothing is broken. Several things produce that variation:
- These systems generate text with a degree of randomness, so wording and the set of brands named can shift between identical prompts.
- The assistant may use live web results in one run and not in another, or retrieve different pages.
- Location, language, account history and personalisation can change what you see.
- Providers update models and retrieval behaviour without notice, so a change in March may be the assistant changing, not your site.
The practical answer is to treat each prompt as a small experiment. Run it several times, then report a rate: “named in two of three runs” is a more honest result than “named” or “not named”. Use a clean browser profile or a consistent setup, and keep it the same every month. I would never present a single screenshot to a client as proof of visibility, and I would not accept one as proof of a problem either. State the limits in your report so nobody mistakes a sample for a census.
Spreadsheet or tool?
There are now many commercial tools that run prompts for you on a schedule. I am not going to compare prices here, because they change and the right choice depends on your size. A decision rule works better than a vendor list.
| Your situation | Sensible choice |
|---|---|
| Fewer than about 50 prompts, one market, monthly review | A spreadsheet and a disciplined hour once a month |
| Many prompts, several languages or locations, weekly reporting | A tool that automates runs, after you have tested it against your manual results |
| You are not sure the exercise is worth it | Do two manual rounds first; the data tells you whether to invest |
If you do buy a tool, ask how it queries the assistants (through an official interface or by simulation), how many runs it averages, and whether you can export the raw answers. A dashboard that shows a “visibility score” without letting you read the underlying answers is hard to trust and impossible to debug.
Cadence and what to report
Monthly is enough for most small and mid-sized businesses. Move to fortnightly while you are making large changes, such as rewriting key pages or launching a new service. Weekly rarely adds information, because the noise between runs is larger than the weekly signal.
Report four things in plain language: your mention rate by bucket, your citation rate and which pages get cited, accuracy problems found, and the competitors and third-party sources that keep appearing. Then pair it with traffic data, since visibility without visits is only half the story. My guide to measuring AI referral traffic in GA4 shows how to see the clicks that do arrive. Together the two give you exposure and response.
What to do with what you find
- Not mentioned at all: assistants lean on what the wider web says about you. Check whether you are described consistently on your own site and elsewhere, and whether credible third parties mention you. My post on unlinked brand mentions covers finding and using them.
- Mentioned but inaccurate: find the page that probably feeds the error. It is often an old directory listing, an outdated press piece or a vague page of yours. Correct the source, and make the right facts easy to quote on your own pages.
- Cited but not named, or named but not cited: look at the page that was used and ask whether it answers the question in its first paragraph.
- Competitors dominate: look at the third-party sources named in the answers. Those are the places the assistant trusts for your category.
I treat my own site the same way. After rebuilding saadrazaseo.com in October 2026 with a custom theme, one connected schema graph, FAQ sections and author information on every article, my check is simply whether assistants describe what I do correctly. I make no claim about results here, because I would rather show you the method than inflate a sample.
Technical access still matters, since a page that cannot be fetched or indexed cannot be cited. If you suspect that, an audit through my technical SEO services is a sensible first step, and the principles in my guide to optimising content for AI Overviews apply to the pages you want cited.
Questions people ask about tracking AI visibility
How many prompts do I need to track?
Start with 30 to 50 that reflect real customer questions, split across category, comparison, problem and brand prompts. That is a working guideline, not an official figure. Choose a number you will still maintain in six months.
Can I rely on one answer from an assistant?
No. The same prompt can produce different answers on different runs, with different brands and sources. Run each prompt several times, report how often you appear, and keep the conditions identical from month to month.
Does Search Console show my AI Overview visibility?
Not separately. Google’s documentation says traffic from AI features is included in the overall Search Console data, so you cannot isolate it there. Manual sampling of the answers is how you see what the feature actually says about you.
Should I pay for an AI visibility tool?
Only after a couple of manual rounds show the work is worth scaling. Pick a tool that lets you read the raw answers and explains how it queries each assistant.
If you want a baseline of where your brand stands, I can build the prompt set and run the first round with you. Contact me for a free audit and we will start from your real customer questions.
More AI search guides: start with the generative engine optimization guide, then go deeper:
- ChatGPT Search SEO: How to Get Your Business Cited
- Perplexity SEO: How to Get Cited in Perplexity Answers
- Google AI Mode SEO: What It Changes and What to Do
- AI Referral Traffic in GA4: How to Measure It Properly
- llms.txt Explained: What It Is and Do You Need One
- AI Crawlers Explained: GPTBot, ClaudeBot, PerplexityBot
- Agentic SEO: How to Make Your Website Work for AI Agents
- Bing SEO for AI Search: Webmaster Tools, IndexNow and Copilot
- Reddit SEO: How Community Content Shapes Google and AI Answers
