Brand monitoring used to mean watching the social web: alerts for your name on X, Reddit, and the news. The question now is different. When someone asks ChatGPT, Perplexity, Gemini, or Google's AI Overviews for the best tool in your category, does the answer say your name? Tracking brand mentions in AI search is how you find out, and it is more measurable than most people assume. This guide covers a free method that actually works, the handful of metrics worth watching, how many times you need to run a prompt before you can trust the result, and the point where a paid tracker starts to pay for itself.
What Tracking Brand Mentions in AI Search Means
A brand mention in AI search is any time an engine names your brand in a generated answer. It comes in three strengths, and the difference matters for what you track:
- Mentioned: the answer names you in the text.
- Cited: the answer names you and links to your site as a source.
- Recommended: the answer puts you forward as the answer, not just a name in a list.
The surfaces worth watching are the ones your buyers actually use: ChatGPT, Perplexity, Gemini, Claude, Google's AI Overviews and AI Mode, and Copilot among others. Each builds its answer a little differently, so a brand mention in one is no guarantee of a mention in another, and your AI visibility has to be read per engine.
Is it even possible to track this reliably? Yes. You are not reading the model's mind, you are sampling its output: you ask the questions your buyers ask, repeatedly, and record whether you show up. The catch is that "repeatedly" is doing real work in that sentence, and it is where most tracking goes wrong.
Why a Single Check Lies
The most common mistake is to ask ChatGPT one question, screenshot the answer, and treat it as the truth. AI answers are not fixed. Ask the same question twice and you can get two different lists of brands, with nothing changed on your side.
Several things move the answer between runs:
- Sampling. The model picks each word from a probability distribution, so the same prompt takes different paths.
- Personalization. Memory, past chats, and account context nudge what you see versus what someone else sees.
- Retrieval. Engines that search the live web pull different pages depending on the moment. ChatGPT may rewrite your question into one or more targeted searches before it answers, and those searches are not identical every time.
- Model routing and prompt wording. A slightly reworded prompt, or a different model version behind the same interface, changes the output.
There is no fixing this; it is how the surface works. A recent study on measuring AI search, Don't Measure Once, puts it plainly: answers vary across runs, prompts, and time, so one-off observations are unreliable, and visibility is best treated as a distribution rather than a single result. The practical consequence is simple. A mention is not a yes or no. It is a rate, and a rate needs more than one measurement.
The Metrics That Matter
Once you accept that you are measuring a rate, the metric set gets clear. You do not need all of these every week, but you should know what each one answers.
| Metric | What it answers | How to read it |
|---|---|---|
| Mention rate | Out of N runs of a prompt, how often are you named? | Your core number. Track it per prompt and per engine, not as one blended figure. |
| Share of voice | Your share of all brand appearances: your mentions divided by every tracked brand's mentions. | Context for the mention rate. Your 30% means more when the strongest rival also sits near 30% than when one competitor is named in almost every answer. |
| Citation rate | How often are you named with a link to your site? | Stronger than a mention: a citation can send referral traffic and shows the answer drew on your page as a source. |
| Position in answer | Are you the recommendation, or the last name in a list? | First-named brands get disproportionate attention. Being present is not the same as being prominent. |
| Sentiment and accuracy | When you are named, is the framing positive, and separately, is it factually right? | Score them as two columns. A confident wrong description is worse than absence, and it is fixable. |
The distinction people trip over most is mention versus citation. A mention is your name in the text; a citation is your name plus a link. Both matter, but they move for different reasons, so track them separately rather than folding them into one score. For a deeper look at the competitive side of this, our guide to AI share of voice walks through how to measure your slice of the answers.
How Many Times to Run a Prompt
This is the question nobody answers, and it is the one that decides whether your tracking means anything. Most guides tell you to "sample more" and leave it there. Here is the actual math, because it changes how you read every number you collect.
A mention rate is a proportion measured from a sample, exactly like a poll. And like a poll, a small sample has a wide margin of error. If you run a prompt 10 times and get named 3 times, your measured rate is 30%, but the true rate could plausibly sit anywhere in a very wide band. That is why a brand can show a 40% rate one week and 25% the next without changing a single page. Nothing changed on the site. The sample was simply too small to separate a real move from noise.
The table below shows roughly how tight your number is at different sample sizes, for a rate near 30% (the margin is widest near 50% and narrower toward the extremes):
| Runs per prompt | Rough 95% margin | What that's good for |
|---|---|---|
| 10 | about ±28 points | A gut check only. "Do I ever show up?" Not a number to report. |
| 30 | about ±16 points | Spotting large gaps against competitors. |
| 100 | about ±9 points | Pinning one period's rate to about ±9 points. |
| ~385 | about ±5 points | Precise benchmarking (this size also covers the worst case near 50%). Rarely worth doing by hand. |
Two rules fall out of this. First, fewer prompts means more runs each - if you only track five prompts, each one carries a lot of weight, so run them more. Second, do not react to a move smaller than your margin. If each week's rate carries a margin of 16 points, a jump from 30% to 38% is not distinguishable from sampling noise. Comparing two periods is actually harder than measuring one, since both readings carry error, so treat small week-over-week moves as inconclusive rather than real. Judging visibility from a handful of prompts checked once or twice is like forecasting the week from one glance out the window.
Two honest caveats. These margins are rough approximations, and at very small samples like 10 runs they are rougher still, so treat the low end as "directional" rather than exact. And runs fired minutes apart share the same live retrieval, so they are not fully independent; spreading runs across a few days gives a truer picture than hammering a prompt in one sitting.
You do not need a statistics degree to apply this. Pick a run count that matches how precise you need to be, keep it consistent, and treat every rate as a range rather than a point.
The Free Way to Track Brand Mentions
You can run a real tracking program with a spreadsheet and an hour a week before you pay for anything. Here is the workflow.
1. Build a prompt set from unbranded questions. The single biggest error is tracking branded prompts like "does Acme have good reviews?" The engine will almost always confirm you exist, which tells you nothing. Track the unbranded questions buyers actually type before they know you: category queries ("best AI visibility tool") and recommendation queries ("what should a small SEO team use to track AI mentions?"). Keep competitor-comparison prompts ("[rival] vs the alternatives") in a separate bucket, since naming a rival measures whether you enter the conversation, which is a different question from unaided category visibility. Aim for 10 to 20 prompts, and put your effort behind the handful that matter most so your headline rate reflects the questions buyers ask most, not an even average across trivia.
2. Run each prompt across the engines, several times. Use fresh or incognito sessions so past chats do not color the answer, and run each prompt the number of times your target precision demands from the section above. Two engines fight this: Google's AI Overviews vary by location and do not fire on every query, and a logged-out ChatGPT can route to a lighter model. Fix your location, note when AI Overviews simply does not appear, and keep your account state consistent so the only thing changing is the model.
3. Log every run in one place. A simple sheet with a row per run: prompt, engine, run number, mentioned (yes/no), cited (yes/no), accurate (yes/no), competitors named, and a note on sentiment. This is the artifact almost every guide skips, and it is the whole game. Without a log you have impressions; with one you have data.
4. Compute your rates. For each prompt and engine, divide mentions by runs to get your mention rate, and count how often each competitor appears for your share of voice. Now you have a baseline you can move.
Be honest about what the free version buys you. By hand you are not going to hit 100 runs across 20 prompts, so do not pretend the spreadsheet gives you week-over-week precision. At 10 to 15 runs on your top 5 to 8 prompts, checked monthly, it answers the questions that matter early: do I show up at all, and who beats me by a wide margin? That is a real, useful read, and it is the honest ceiling of manual tracking. The paid tools exist to buy back the precision and the hours, not to make the free method fake.

To speed up the checking itself, a free scanner does the multi-engine legwork for you. Our free AI Visibility Checker runs one prompt across the eight major engines in a single pass and marks whether each one cites you, mentions you, or recommends you, with a free export you can paste straight into your log. It is a single scan rather than continuous monitoring, which is exactly what you want when you are still building the habit for free.
Track Every Engine, Not Just ChatGPT
Checking only ChatGPT is the second big blind spot. The engines disagree, often sharply. You can be the top recommendation in one and absent from another for the exact same question, because each pulls from different sources and, for the ones that search live, assembles the answer from different retrieved pages.
Google's AI Overviews, for example, may break a single question into several sub-queries through a fan-out step, then draw citations from the organic index. ChatGPT may use its own search partners and rewrite your query before answering. Perplexity, Gemini, and Claude each have their own retrieval and their own preferences. The result is that a mention rate is only meaningful per engine, and one engine's picture can quietly mislead you about the others. Our breakdown of how ChatGPT cites sources shows why the same brand can win one engine and lose the next.
You do not have to track all of them equally. Weight your effort toward the engines your buyers actually use, and check the rest less often. But check them. A program that watches one engine is measuring one slice and calling it the whole.
When a Tracking Tool Is Worth It
The free method has a ceiling. Twenty prompts, across five engines, run thirty times each, week after week during an active push, is three thousand checks logged by hand. That is the point where a tool stops being a luxury and starts saving real time.
The trap to avoid is the dashboard that shows you the problem and stops there. Plenty of tools will tell you your mention rate is low and leave you to figure out what to do about it, which is why marketers have started to question whether expensive AI visibility tools earn their price. What a good tool actually needs: continuous tracking across the engines, share of voice against named competitors, a clear week-over-week delta so you can see moves against the noise, and a path from the number to the fix.
It should also help you close the loop on traffic. AI answers that link to you send referral visits, and you can see them in GA4 by watching for referrers like chatgpt.com, perplexity.ai, and gemini.google.com. Treat that traffic as the visible tip only: most AI answers resolve without a click, so referral sessions undercount how often you actually influence a buyer. That undercount is the reason to track the mention rate directly rather than inferring your AI presence from clicks alone.
Across the sites we scan for AI visibility, the most common problem is not the score itself. It is that teams are flying blind: a single engine, a handful of branded prompts, checked once, mistaken for the full read. That is the gap our AI Brand Monitoring dashboard is built to close, tracking citation share and share of voice across all eight engines over time with a week-over-week delta. Like any tracker, it samples rather than delivering a fixed verdict, so the same run-count discipline applies. If you would rather compare the field first, our roundup of the best AI visibility tools lays out the options.
What to Do When You Find a Gap
Tracking is only half the work. The question that follows every low mention rate is the same: the engine recommends my competitor and not me, so why, and what now?
The why is usually not a secret algorithm. Generative engines build answers from the third-party sources they trust and retrieve, and the research behind generative engine optimization found that how content is written and cited materially changes whether an engine surfaces it. Your competitor is named clearly on the pages the engine pulls from, and you are missing, thin, or described with stale facts. For the engines that retrieve live sources, the answer largely repeats what those sources say rather than judging your product directly; base-model answers move more slowly but still lean on how widely and consistently you appear in training data.
So the fix follows the log. Take the prompts you lose and look at which sources the engine cites in those answers. Then go earn a presence there: the review sites, the listicles, the community threads, and the industry media that keep appearing. One mention rarely moves anything; what shifts an answer is consensus, the same brand described consistently across several independent sources the engine already trusts. Make sure your own facts line up across the web so the engine has a clear, repeated signal, which is the core of strong entity SEO. Then re-track the same prompts and watch the rate settle over the next few weeks.
A worked example: say your mention rate for "best AI mention tracker" is 10% in Perplexity, and the answers keep citing the same three comparison articles. You do not have to move the model. You have to get accurately represented in those three articles, then confirm the lift over your next measurement window. That is the whole loop, and it is why tracking and fixing belong in the same workflow rather than two separate tools. For a structured version of this, our AI visibility audit walks through it step by step.
Frequently Asked Questions
Is it possible to track brand mentions in AI search?
Yes. You cannot see inside the model, but you can sample its output: run the prompts your buyers ask, repeatedly, across each engine, and record whether your brand appears. Because answers vary run to run, you track a mention rate over many runs rather than reading a single response. That is measurable with a spreadsheet or a tool.
How often should I check whether AI engines mention my brand?
Match the cadence to how fast your category moves and how much you are actively working on it. Monthly is enough for a stable baseline; weekly makes sense when you are running a campaign to earn mentions. More important than frequency is consistency: same prompts, same run count, so the numbers are comparable over time.
What's the difference between a brand mention and a citation in AI answers?
A mention is your name in the generated text. A citation is your name plus a link to your site as a source. A citation is the stronger signal because it can send referral traffic and shows the answer drew on your page as a source. Track them as two separate metrics, since they move for different reasons.
Are there free tools to track brand mentions in AI search?
Yes, for the checking step. Free scanners, including our own AI Visibility Checker, run a prompt across the major engines and show whether you are cited, mentioned, or recommended. They are single scans rather than continuous monitors, so pair them with a logging spreadsheet to build a rate over time. Ongoing, automated tracking across many prompts is where paid tools come in.
Why do the AI answers change every time I check?
Because the models sample their output and, for the ones that search live, retrieve different pages each time. Personalization, prompt wording, and model routing add more variation. This is normal and unavoidable, which is exactly why a single check is unreliable and you measure a rate across repeated runs instead.
Is tracking ChatGPT enough, or do I need the other engines?
Not enough on its own. Engines pull from different sources and can disagree sharply, so you can lead in ChatGPT and be absent in Gemini for the same question. Track the engines your buyers actually use, and check the others less often rather than ignoring them.
Start Tracking This Week
The reason to start now is not that AI search is coming, it is that it already answers without a click. Bain reports that about 60% of searches now end without the user moving on to another site, a shift it ties to the rise of AI summaries. When the answer is the destination, whether it names you is what you need to know. You do not need budget to find out where you stand. Build an unbranded prompt set, run it across the engines a few times, and log what comes back. Run a first pass through the free AI Visibility Checker to see your eight-engine picture in minutes, and when the spreadsheet gets heavier than the insight, move the whole thing into the AI Brand Monitoring dashboard and let it keep the rate for you.
Sources
- Goodbye Clicks, Hello AI: Zero-Click Search Redefines Marketing - Bain & Company, 2025 -
bain.com/insights/goodbye-clicks-hello-ai-zero-click-search-redefines-marketing/ - Don't Measure Once: Measuring Visibility in AI Search (GEO) - arXiv, Schulte, Bleeker & Kaufmann, April 2026 -
arxiv.org/abs/2604.07585 - ChatGPT Search - OpenAI Help Center -
help.openai.com/en/articles/9237897-chatgpt-search - AI Features and Your Website - Google Search Central -
developers.google.com/search/docs/appearance/ai-features - GEO: Generative Engine Optimization - arXiv, Aggarwal et al. (ACM SIGKDD 2024) -
arxiv.org/abs/2311.09735 - Marketers question expensive AI visibility tools as inconsistent results fuel skepticism - Digiday, Kimeko McCoy, May 2026 -
digiday.com/marketing/marketers-question-expensive-ai-visibility-tools-as-inconsistent-results-fuel-skepticism/