AI brand monitoring guide showing mention rate by engine across Perplexity, AI Overviews, Gemini, Copilot and ChatGPT with cost per tracked AI answer

Somewhere today, a buyer asked ChatGPT which vendor they should shortlist in your category. You will never see that query. It will not appear in Google Search Console, it will not show up in your analytics, and if the answer named three competitors and not you, nothing anywhere will tell you it happened.

That is the gap AI brand monitoring exists to close. Done properly it tells you how often AI engines name you, where you sit in the answer, whether what they say is even true, and which pages they read to decide. Done badly, which is how most of it is done, it produces a number that feels precise and means almost nothing.

This guide covers what to track, the sampling problem that quietly invalidates most checks, a free method you can run in half an hour, and every major tool compared on pricing we verified from the vendors' own pages in September 2026. It also includes a cost-per-answer analysis you will not find elsewhere, because it is the only way to compare these tools honestly.

What AI brand monitoring actually tracks

Traditional brand monitoring watches for your name in places that publish: news, blogs, forums, social. AI brand monitoring watches something structurally different. It watches what a model says when a person asks it a question, which is not published anywhere and does not exist until the moment it is generated.

That difference matters more than it sounds. A press mention is a fixed artifact. You can link to it, screenshot it, count it. An AI answer is generated fresh each time, varies between users, and disappears the moment the chat closes. You are not monitoring a corpus. You are sampling a probability distribution.

Everything difficult about this category follows from that one fact.

This is not social listening with a new label

Plenty of vendors have bolted the words "AI" onto an existing mention-tracking product. It is worth being clear about why that does not work.

Social listening answers "who talked about us?" It queries an index of things humans posted. The ground truth exists independently of the tool, and two tools querying the same index should broadly agree.

AI brand monitoring answers "what would a model tell a buyer?" There is no index to query. The tool has to actually ask the model, repeatedly, and record what comes back. Two tools asking the same question at the same moment can legitimately get different answers, and neither is wrong.

This is why a monitoring tool's value is almost entirely in its sampling design, not its dashboard. The dashboards all look similar. What separates them is how many questions they ask, how many models they ask, and how often.

The four things worth measuring

Most dashboards in this category present a single composite "visibility score". Those scores are not comparable between tools, because every vendor weights the inputs differently, and a score that cannot be compared or reproduced is a vanity metric. Track the four underlying numbers instead.

What to trackThe question it answersWhy it matters
Mention rateOut of the buying prompts your customers ask, how many name you at all?The headline number, and the one that moves first when anything improves.
Position in the answerAre you named first, or sixth of eight?Readers rarely reach the bottom of a generated list. Position is most of the value.
Accuracy and sentimentIs what the model says about you actually correct?A wrong price or a dead feature talks a buyer out of the demo. Worse than silence.
Cited sourcesWhich pages did the engine read to produce that answer?The only column that tells you what to actually go and fix.

If you only ever look at one, look at cited sources. Mention rate tells you that you have a problem. Cited sources tell you where it lives.

The sampling problem nobody warns you about

Here is the failure that invalidates most of the monitoring people do.

You ask ChatGPT "what are the best tools in [your category]" once. You are named second. You screenshot it, you feel good, you move on. A week later a colleague runs what they believe is the same check and you are absent entirely. One of you concludes the tool is broken. Neither of you is wrong.

Generated answers vary for reasons that have nothing to do with your marketing:

  • The model samples probabilistically, so identical prompts return different text
  • Account history and memory personalise what you specifically see
  • Retrieval is live, so the pages pulled in differ between sessions
  • Models are updated continuously, without announcement
  • Phrasing changes the answer, and so does the country you ask from

A single check is one draw from a distribution. Treating it as a measurement is the single most common mistake in this category, and it is the reason two people at the same company routinely disagree about whether AI "knows" them.

The fix is not a better tool. It is more draws.

How much sampling is actually enough

You do not need academic rigour here, but you do need enough repetition that a change in the number means something. A practical floor for a monthly read:

  • At least 20 prompts, covering the different ways buyers actually ask, not 20 rewordings of one question
  • At least 3 engines, because coverage differs sharply between them
  • At least 3 runs per prompt, in a clean session, spread across days rather than fired in one sitting
  • Recorded, not remembered, so you are comparing to a written baseline rather than an impression

Twenty prompts across three engines run three times is 180 observations a month. That is enough to notice a real shift and to stop chasing noise. It is also, not coincidentally, roughly where paid tools start to earn their price, because doing it by hand at that volume is genuinely tedious.

One discipline matters more than volume: always check in a clean session. Logged in, with history and memory active, a model that has watched you research your own company for months will name you far more readily than it names you to a stranger. That is the single most flattering and most useless reading you can take.

Which engines to monitor, and the one everybody forgets

Coverage is not interchangeable. Each engine builds answers differently, which is covered in depth in our breakdown of how the major engines choose which brands to mention. For monitoring purposes the practical differences are these.

EngineHow it builds the answerWhat that means for monitoring
ChatGPTSynthesises heavily, cites sparingly, leans on broad consensus across the web.Highest personalisation risk. Always check in a temporary chat.
PerplexityCitation-first, surfaces sources inline for almost every claim.The most useful engine to monitor, because it shows which pages to fix.
Google AI OverviewsTied to Google's index, ranking signals and structured data.Closest to classic SEO. Often the fastest surface to move if you already rank.
GeminiGoogle ecosystem signals, strongly multimodal, rewards freshness.Frequently diverges from AI Overviews despite sharing an index. Monitor both.
ClaudeFavours clear, reasoned, low-noise explanation over optimised copy.Rewards genuine depth. Over-optimised pages tend to do worse here.
Microsoft CopilotSits inside Microsoft 365, weights integration and enterprise signals.The one everybody forgets. Your enterprise buyer lives here all day.

Copilot is the consistent blind spot. It is the engine embedded in the software your enterprise buyer already has open, and it is the one least likely to appear in a monitoring setup, partly because several tools charge extra for it or omit it entirely.

What to do when the engines disagree

They will, and often dramatically. You will be named in 40 percent of Perplexity answers and 5 percent of ChatGPT answers on the identical prompt set. New teams treat this as a data problem and go looking for the broken engine. It is not a data problem. It is the finding.

Disagreement is diagnostic, and the pattern tells you where the weakness is.

Strong in Perplexity, weak in ChatGPT. Perplexity retrieves live and cites as it goes, so it rewards having relevant, findable pages right now. ChatGPT leans more on accumulated consensus. This pattern usually means your content is good but your brand is not yet widely enough described for a model to name you from general knowledge. Keep going, it is the normal shape for a company building authority.

Strong in ChatGPT, weak in Perplexity. The reverse, and rarer. It usually means brand familiarity exists but the specific pages that would support a cited answer are missing, thin or hard to retrieve. This is the more fixable of the two.

Strong in AI Overviews, weak everywhere else. You have good SEO and little else. Your rankings carry you on Google surfaces, but nothing outside Google's index describes you well enough for the other engines. Very common in companies with a long-standing SEO programme and no third-party footprint.

Strong everywhere except Copilot. Often an enterprise-signal gap: thin integration documentation, no presence in Microsoft-adjacent sources. Worth fixing only if you sell into enterprise IT, in which case it is worth fixing urgently.

Record the per-engine split rather than averaging it away. A single blended score hides exactly the information that tells you what to do on Monday.

How to run AI brand monitoring without buying anything

Before you spend money, run this once. It takes about thirty minutes and it will tell you whether you have a problem worth paying to solve.

1. Write ten buying prompts. Not "what is [your company]". Nobody who does not know you types that. Write what a buyer who has never heard of you would ask. "Best [category] for [company size]". "Alternatives to [your biggest competitor]". "Which [category] tool integrates with [the platform your ICP runs on]". "[Competitor] vs [competitor]".

2. Open a clean session in each engine. Temporary chat in ChatGPT, a logged-out or private window elsewhere. This step is not optional and it is the one people skip.

3. Run every prompt in each engine and record four columns. Were you named, in what position, was anything said about you wrong, and which sources were cited. A spreadsheet is entirely sufficient.

4. Repeat the whole set twice more over the following week. One pass is an anecdote. Three passes is a baseline.

5. Count. Mentions divided by total observations is your mention rate. Write it down with the date. That number is now the thing you are trying to move.

If you would rather not do this by hand for a first read, the free AI Visibility Checker runs a version of this across engines and returns a report in a couple of minutes. It will not replace a disciplined monthly cadence, but it is a faster way to find out whether the number is 60 percent or zero.

Building a prompt set that reflects real buying

The quality of your monitoring is capped by the quality of your prompt list. Most lists fail the same way: they are written by someone who already knows the product, so they accidentally describe it.

A prompt set that reflects real demand covers five shapes:

  • Category discovery. "Best tools for [job to be done]" with no vendor named
  • Head-to-head. "[Competitor A] vs [Competitor B]", including pairs you are not in
  • Switching. "Alternatives to [competitor]", which is pure displacement intent
  • Constraint-led. "[Category] tool that works with [stack]" or "under [budget]"
  • Vertical. "[Category] for [industry]", where specificity beats size

Write them the way a buyer types, in full sentences, with the messy detail included. Keep them fixed once set. Changing prompts month to month makes your trend line meaningless, and the trend is the entire point.

Should you buy a tool?

Be honest about what you are buying. You are not buying access to data nobody else can get. Anyone can open ChatGPT. You are buying three things: scheduling, so the checks happen without a human remembering; scale, so it is 500 observations a month rather than 30; and history, so you have a trend line instead of a folder of screenshots.

Those are real. They are also worth different amounts to different teams.

Do it manually if you are testing whether the channel matters, if you sell in one category with one clear competitor set, or if you can genuinely commit thirty minutes a month.

Buy a tool if you need to report a trend to someone who was not in the room, if you are tracking multiple brands, regions or languages, or if you have already found a problem and need to prove the fix worked.

AI brand monitoring tools compared, with verified pricing

Every figure below was read from the vendor's own pricing page in September 2026. Where a vendor does not publish pricing, we say so rather than guessing.

ToolEntry pricePrompts at entryEngines at entryFree option
Arobis AI Visibility CheckerFreeFixed scanMulti-engineYes, no card
AthenaHQ EssentialFree300 credits5 enginesYes, $25 credit
Otterly.ai Lite$29/mo154 engines, Claude and Gemini cost extraNo
HubSpot AEO$50/mo25 cap3, no Google coverageNo
Peec AI Starter$95/mo50Choose 3 modelsNo
Scrunch Core$250/mo1254 engines7-day trial
AthenaHQ Starter$295/mo3,600 credits11 modelsEssential tier
ProfoundNot publishedNot publishedBroadNo

A few things in that table are worth saying out loud.

AthenaHQ has a genuinely free tier, with $25 of credit and five engines included. It is the most generous free entry point of any paid platform here, and it is routinely left out of comparison posts.

Otterly's $29 plan covers four engines but not Claude, Gemini or Google AI Mode, which are sold as add-ons. If your buyers use Gemini, the real entry price is higher than the headline.

Scrunch now sits at $250 per month for Core with a 7-day free trial, and the company was acquired by Sitecore in 2026. Some comparisons, including earlier versions of ours, quoted $300. The current published Core price is $250.

Profound does not publish pricing at all. Treat any specific figure you read for Profound, anywhere, as a buyer report rather than a quoted price.

If you want the long-form teardown of any single platform, we maintain detailed comparisons of the best Peec AI alternatives, the best Scrunch AI alternatives, the best AthenaHQ alternatives, the best Otterly AI alternatives and the best Profound alternatives, plus a full review of HubSpot's AEO grader after 28 days of use.

How to read a pricing page in this category

Headline prices in this market are close to meaningless, because the unit being sold is not the same from vendor to vendor. One sells prompts, one sells credits, one sells projects. The only number that compares is the cost of one tracked AI answer.

The formula every vendor is implicitly using is the same:

prompts x models x checks per month = AI answers per month

Peec AI states this outright on its pricing page. Applying it to the published entry plans, using daily tracking where the vendor states daily tracking:

PlanThe mathsAnswers per monthCost per answer
Otterly.ai Lite, $2915 prompts x 4 engines x 30 days1,800About $0.016
Otterly.ai Standard, $189100 prompts x 4 engines x 30 days12,000About $0.016
Peec AI Starter, $9550 prompts x 3 models x 30 days4,500About $0.021
Peec AI Pro, $245150 prompts x 3 models x 30 days13,500About $0.018
AthenaHQ Starter, $2953,600 credits, 1 credit = 1 AI response3,600About $0.082
Scrunch Core, $250125 prompts x 4 LLMs, frequency not publishedCannot be calculatedAsk before buying

That table reframes the market. On a cost-per-answer basis the cheap tools are not slightly cheaper, they are roughly five times cheaper. AthenaHQ's Starter plan is not badly priced for what it is, but you are paying for eleven-model breadth and an agent layer, not for volume.

The credit model deserves particular attention, because it is easy to misread. AthenaHQ's 3,600 monthly credits sound generous. Spread across daily tracking they work out as:

  • Across 3 models: about 40 prompts trackable daily
  • Across 5 models: about 24 prompts trackable daily
  • Across all 11 models: about 11 prompts trackable daily

In other words, the breadth that justifies the price is the same thing that consumes it. Use every model and $295 a month buys you eleven questions a day. That is not a criticism of the product, it is an argument for deciding your prompt set and engine list before you choose a plan, because the two decisions are really one decision.

Ask every vendor the same three questions before you pay: how often is each prompt refreshed, does a failed or empty response consume quota, and what happens when you run out mid-month.

Red flags when evaluating a vendor

Demos in this category look identical. Every dashboard is a line going up and to the right with your logo next to a competitor's. Use these to tell them apart.

A headline score with no published method. If the vendor cannot tell you exactly how the number is calculated, how many runs it averages and over what window, it cannot be audited or reproduced. Ask them to explain the formula. A good vendor answers immediately.

No per-engine breakdown. A blended figure across engines hides the diagnostic pattern described above, which is the most useful thing in the whole data set. If you cannot split it by engine, walk away.

Sampling frequency that is vague or absent. "Continuous monitoring" is marketing. Ask how many times a specific prompt is run per week and what the answer is averaged over. If a prompt is checked once a week, your monthly number rests on four observations.

No raw answers. You should be able to read the actual text the model produced, not just a tally. Without the raw output you cannot check accuracy, and accuracy is the finding most likely to be costing you deals right now.

Cited sources missing or shallow. If the tool tells you that you were mentioned but not what the engine read, it has given you a symptom and withheld the diagnosis.

Prompts you cannot edit. Some tools generate a prompt set for you and lock it. Your prompt set is the entire measurement instrument. You need to own it.

Annual contract before a trial. This market is a year old and moving fast. Anyone requiring twelve months up front, before you have seen your own data, is managing their churn risk with your budget.

What monitoring cannot tell you

This is the part vendors are quiet about, and it is the most important section here.

Monitoring is a thermometer. It tells you the temperature. It does not tell you why the room is cold and it does not heat the room.

A dashboard showing 12 percent mention rate is a fact, not a diagnosis. It cannot tell you whether you are absent because your category positioning is unclear, because no third party has ever written about you, because your pricing page is gated so models cannot read it, or because a competitor simply has ten years of reviews you do not.

Worse, mention rate is not directly actionable. You cannot go and do "mention rate" on Monday morning. The number that is actionable is the one in the cited sources column, because it points at specific pages on specific domains, and those can be fixed, earned or corrected.

We have written separately about the difference between watching the number and moving it, which is the distinction most teams discover about three months into a monitoring subscription.

From monitoring to moving the number

Once you have a baseline and a source list, the work divides into three buckets, roughly in order of how fast they pay back.

Fix what the models can already see. If an engine is reading your product page and getting your pricing wrong, that is a content problem you control entirely. Make the facts explicit, current and machine-readable. Publish an llms.txt file so models have a clean summary of what you do. This is the cheapest work in the list and often the fastest to move.

Earn what the models trust. The cited sources column will keep pointing at the same handful of third-party domains: review platforms, comparison articles, community threads. Those are the pages actually deciding the answer. Being present and accurate there matters more than another post on your own blog.

Own the category language. Models answer the question as asked. If buyers ask in vocabulary your site never uses, you are invisible regardless of authority. This is the failure behind a surprising share of zero scores.

For a sense of what good looks like across a market rather than a single brand, our State of AI Search Visibility research sets out the benchmarks.

Connecting monitoring to something a CFO recognises

Mention rate will not survive a budget conversation on its own. It is a new metric with no history and no obvious link to revenue, which makes it easy to cut.

Three framings that do survive that conversation:

  • Share of the shortlist. Of the buying prompts in your category, what percentage produce an answer naming you? That is directly comparable to share of voice, which finance already understands.
  • Competitive displacement. Track the same prompts for your three main competitors. "We are named in 12 percent, they are named in 61 percent" is a number that gets attention in a way an absolute score never does.
  • Correctable error rate. How many answers contain something factually wrong about you? This reframes the spend as risk management, which is an easier budget to defend than awareness.

Pair the number with pipeline language. If your sales team can point to one deal where a buyer arrived having already been told something wrong by an assistant, that single anecdote does more than any dashboard.

Who owns this, and what it actually costs in time

Monitoring programmes usually die for organisational reasons rather than technical ones. Nobody owns the number, so nobody looks at it, and the subscription gets cancelled at renewal.

In practice the work splits cleanly.

Someone owns the number. Usually whoever owns organic search or demand generation. One named person, not a team. Their job is to run the check, record it and report the trend.

Someone owns the content fixes. Product marketing, usually, because most fixes are positioning and factual accuracy rather than writing. Wrong pricing, missing integration details, category language that does not match how buyers ask.

Someone owns the third-party surface. Review profiles, directory listings, comparison pages you do not control. This frequently belongs to nobody, which is precisely why it is the weakest signal for most companies.

The realistic time cost, once set up: about two hours a month to run a twenty-prompt set across three engines manually and record it, or about thirty minutes a month if a tool runs the sampling and you only interpret the output. Budget another half day per quarter for the report that goes to leadership.

That is the whole commitment. Programmes that fail rarely fail because two hours was too much. They fail because the two hours belonged to nobody in particular.

A realistic 90-day plan

Days 1 to 7. Write the prompt set. Run it manually across three engines in clean sessions. Record the four columns. This is your baseline and you will compare against it for a year, so write it down properly.

Days 8 to 30. Fix what you control. Correct wrong facts on your own pages, make pricing and integration details readable, publish llms.txt. Do not expect the number to move yet.

Days 31 to 60. Work the cited sources list. Get accurate on the third-party pages the engines keep reading. Run the prompt set again, same prompts, same method, and compare honestly.

Days 61 to 90. Decide on tooling with evidence rather than optimism. You now know your prompt count, your engine list and whether the number moves at all, which is exactly what you need to choose a plan without overbuying.

If at day 90 the number has not moved, the constraint is almost certainly authority rather than content, and that is a different and slower project. Our work on AI search demand generation for B2B SaaS covers what that looks like, and our plans and pricing are public if you would rather hand it over.

Frequently asked questions

What is AI brand monitoring?

AI brand monitoring is the practice of systematically tracking how AI assistants such as ChatGPT, Perplexity, Gemini, Claude and Copilot describe and recommend your brand. It measures how often you are named in answers to buying questions, where you appear in those answers, whether the information is accurate, and which sources the engine used.

How is AI brand monitoring different from traditional brand monitoring?

Traditional monitoring searches an index of things people published. AI brand monitoring samples answers that are generated fresh each time and exist nowhere until they are produced. Because the output varies between runs and between users, it has to be sampled repeatedly rather than looked up once.

How often should I check my AI brand visibility?

Monthly is enough for most B2B companies, provided each check runs the same fixed prompt set several times across several engines. Weekly is worth it in fast-moving categories or while you are actively working on a fix. Checking once and treating it as a measurement is the most common error.

Do I need a paid tool for AI brand monitoring?

No. A spreadsheet, a fixed prompt set and thirty minutes a month will give you a defensible baseline. Paid tools buy scheduling, scale and a trend line rather than access to data you could not otherwise get. Buy one once you know your prompt count and engine list, not before.

How much does AI brand monitoring cost?

Verified entry prices in September 2026 run from free, through Otterly.ai at $29 per month for 15 prompts and Peec AI at $95 per month for 50 prompts, to Scrunch at $250 and AthenaHQ at $295 per month. AthenaHQ also offers a free Essential tier. Profound does not publish pricing.

Why do AI answers about my brand keep changing?

Because generation is probabilistic, retrieval is live, and models are updated continuously. Your account history also personalises what you see. Two people asking the same question at the same moment can legitimately get different answers, which is why single checks are unreliable.

Why does AI recommend my company to me but not to my customers?

Almost always because you are checking while logged in to an account that has spent months reading about your company. Memory and history make the model far more likely to name you. Always check in a temporary or logged-out session.

Can AI brand monitoring fix bad information about my company?

Not by itself. Monitoring finds the error and shows you which source produced it. Correcting it means fixing the underlying page, whether that is your own site or a third-party listing, and then waiting for the engines to re-read it.

Which AI engines should I monitor first?

Begin with ChatGPT for reach, Perplexity because it shows its sources, and Google AI Overviews because it is closest to the SEO you already do. Add Microsoft Copilot early if you sell into enterprise IT, since your buyers spend the day inside Microsoft 365.

What is a good mention rate?

There is no universal benchmark, because it depends entirely on category competitiveness. The useful comparison is not against an absolute number but against your own baseline over time and against the competitors you track with the identical prompt set.

Does AI brand monitoring replace SEO?

No. Google AI Overviews and Gemini lean heavily on the same signals classic SEO produces, so strong SEO usually helps. But ranking first on Google does not guarantee being named in a generated answer, which is why the two need measuring separately.

How long before the number moves?

Fixing factual errors on pages engines already read can show up within weeks. Building the third-party authority that changes whether you are recommended at all is a matter of months, not weeks. Anyone promising faster is selling something.

Why do different monitoring tools give me different numbers?

Because they ask different questions, different numbers of times, across different engines, and average them differently. Two tools can both be correct and disagree. This is why the underlying method matters more than the headline score, and why switching tools resets your trend line.

The short version

AI brand monitoring is worth doing, and most of it is being done wrong. The errors are consistent: checking once instead of sampling, checking while logged in, tracking a composite score instead of cited sources, and buying a tool before knowing what needs tracking.

Get the method right and it costs nothing to start. Run ten real buying prompts across three engines in clean sessions, three times, and write down what you find. That single exercise will tell you more than most dashboards, and it will tell you whether the number is worth paying to watch.

Then remember what the number is for. Monitoring tells you the temperature. Changing it is a different job.

Keep Reading