Type your category into ChatGPT and ask which tools it recommends. If your company is named, you feel good for about a minute. If it is not, you feel something closer to panic.
Both reactions are usually based on a broken test. The single most common way people check their ChatGPT visibility produces a result that is systematically wrong, and wrong in the direction that flatters you. Most brands that believe they are visible in ChatGPT are not.
This guide explains what ChatGPT visibility actually is, why your own check is probably lying to you, how to run one that is not, what genuinely decides whether ChatGPT names a company, and what to do when the answer is zero.
What ChatGPT visibility actually means
ChatGPT visibility is how often ChatGPT names your brand when someone asks a question your product could answer, and how it describes you when it does.
It is not a ranking. There is no position one. There is no ChatGPT equivalent of a search results page where ten links sit in order and you occupy a slot. There is a generated paragraph, it mentions somewhere between zero and six companies, and you are either in it or you are not.
That binary is what makes this different from search, and harder. In Google you can rank eleventh and still get traffic, still be findable, still be in the running. In a generated answer, eleventh does not exist. Being left out is not a weak position. It is absence.
Three things are worth separating, because people collapse them and then confuse themselves:
- Brand knowledge. Does ChatGPT know who you are if asked directly by name?
- Recommendation. Does it bring you up unprompted when someone describes a problem you solve?
- Accuracy. When it does talk about you, is what it says actually true?
Almost every company has more brand knowledge than recommendation. ChatGPT can often produce a paragraph about you when you name yourself, and still never mention you to a buyer who does not know you exist. Recommendation is the one worth money, and it is the one people forget to test.
Why your own check is almost certainly lying to you
Here is the thing that invalidates most self-assessments in this category.
When you ask ChatGPT about your own market, you are almost certainly signed in to an account that has spent months or years reading, writing and researching your company. That account has memory. It has a history full of your product, your competitors, your positioning documents, your launch copy.
ChatGPT uses that. It is designed to. Personalisation is a feature, and in this one narrow case it is a trap.
The result is that the model is dramatically more likely to name your company to you than it is to name it to a stranger in another country who has never typed your name. You are not measuring your visibility. You are measuring the model's memory of your own browsing.
We have checked this pattern repeatedly and the gap is not subtle. A brand that appears consistently in a logged-in account can disappear entirely in a clean session run seconds later on the same prompt.
If you take one thing from this article, take this: a ChatGPT visibility check performed while logged in to your working account is worthless. It is the equivalent of asking your mother whether your business idea is good.
The clean-session test
The fix takes about fifteen seconds to set up.
In ChatGPT, start a temporary chat. It is in the model or account menu depending on your interface version, and it means the conversation has no memory, does not use your history, and is not saved. This is the closest you can get to seeing what a stranger sees.
Belt and braces, if you want to be thorough:
- Use a temporary chat rather than just a new chat, because a new chat still draws on saved memory
- Or open a private or incognito window and use ChatGPT logged out entirely
- Or ask a colleague in another country to run the identical prompt and send you the raw answer
- Do not include your brand name in the prompt unless you are explicitly testing brand knowledge
Run the same prompt in both modes, your normal account and a temporary chat, and compare. That side-by-side is often the most persuasive internal argument you will ever make about this topic, because leadership can see the difference on one screen.
How to check your ChatGPT visibility in about ten minutes
Step 1. Write eight to ten buying prompts. These must be questions asked by someone who does not know you. Good shapes:
- "Best [category] tools for [company type]"
- "Alternatives to [your largest competitor]"
- "What should I use for [the job your product does]"
- "[Competitor A] vs [Competitor B]", including matchups you are not part of
- "[Category] software that integrates with [platform your buyers run]"
- "Cheapest [category] tool for a small team"
Step 2. Open a temporary chat. Non-negotiable, for the reasons above.
Step 3. Run each prompt and record four things. Were you named at all. In what position in the list. Was anything said about you incorrect. Which sources, if any, were referenced.
Step 4. Run the whole set again twice more, on different days. Generated answers vary between runs even with identical inputs. One pass is an anecdote, three is a baseline.
Step 5. Calculate. Mentions divided by total observations. Ten prompts run three times is thirty observations. Named in six of them is a 20 percent mention rate. Write the number and the date down somewhere permanent.
That number is your baseline. Everything you do from here is measured against it, which is why the method has to stay identical each time you repeat it.
If you want a faster first read without the manual work, the free AI Visibility Checker runs a multi-engine scan and returns a report in a couple of minutes, and there is a dedicated ChatGPT visibility checker if ChatGPT is the only engine you care about today.
A prompt set you can copy
The most common reason a first check produces a useless result is a weak prompt list. People write prompts that describe their product, because they know their product. Buyers do not do that.
Replace the bracketed parts below and you have a working twenty-prompt set. Keep them in a document and never change them, because the trend is what matters.
Discovery, where no vendor is named
- What are the best [category] tools right now?
- What [category] software do [your ICP, for example B2B SaaS companies] use?
- I need to [job your product does]. What should I use?
- What is the best [category] tool for a team of [size]?
- Which [category] platform is easiest to set up?
Switching, which is pure displacement intent
- What are the best alternatives to [largest competitor]?
- We are unhappy with [competitor]. What else should we look at?
- Is there a cheaper alternative to [competitor]?
- What is a good [competitor] replacement for a smaller team?
Head to head, including matchups you are not in
- [Competitor A] vs [Competitor B], which is better for [use case]?
- How does [competitor] compare to other [category] tools?
- Who are [competitor]'s main competitors?
Constraint led, where specificity beats size
- Which [category] tool integrates with [platform your buyers run]?
- Best [category] software under [budget] per month
- [Category] tool with an API and SSO
- Which [category] tools are GDPR compliant and EU hosted?
Vertical, where you can beat far larger competitors
- Best [category] software for [industry]
- What do [role, for example heads of demand generation] use for [job]?
- [Category] tool for regulated industries
- Best [category] platform for a company scaling from [X] to [Y]
Notice that your brand name appears nowhere. That is deliberate. Including it tests whether ChatGPT knows you, which is a much easier test and a much less useful one.
What actually drives whether ChatGPT names you
ChatGPT is not consulting a ranking table. It is producing the most plausible answer it can, and plausibility here is built from repetition and consensus across everything it has read.
In practice, four things decide it.
Being described consistently, by other people. The strongest signal is not what your website says about you. It is whether many independent sources describe you the same way. Review platforms, comparison articles, community threads, podcasts, directories. A model that has read fifty sources calling you "a [category] tool for [audience]" will produce that sentence. A model that has only read your own homepage will hesitate.
This is the single biggest lever, and it is also the slowest, which is why people avoid it in favour of things that feel faster.
Category clarity. Models answer the question as asked. If a buyer asks for a "[category] tool" and your site never uses that phrase, preferring language you invented, you will not be surfaced no matter how good you are. Inventing category language is a positioning choice with a direct visibility cost.
Machine-readable substance. Pricing behind a form, features described only in a video, key facts trapped in an image. If it cannot be read, it cannot be repeated. Publishing an llms.txt file is a cheap way to hand models a clean, correct summary of who you are.
Corroboration over assertion. Your own claims count for little. "Best-in-class" on your homepage is noise. The same claim on three independent sites is signal. This is why review profiles matter so much more than another blog post.
The deeper mechanics, including how this differs across engines, are in our breakdown of how the major engines choose which brands to mention.
Where ChatGPT actually gets its answers
It helps to know which surfaces are doing the work, because that is where the fixing happens.
| Source type | Why it carries weight | What you can do about it |
|---|---|---|
| Review platforms | Structured, comparative, and written by third parties rather than by you. | Keep profiles complete and current. Volume and recency both matter. |
| Comparison and listicle pages | Explicitly name vendors side by side, which is the shape of a buying answer. | Be present and described accurately, including on pages you do not control. |
| Community discussion | Reads as unpaid opinion, which models treat as strong corroboration. | Earn it. It cannot be bought convincingly and attempts to fake it age badly. |
| Your own product pages | Treated as the authority on facts: pricing, features, integrations. | Make them explicit, current and readable without a form. |
| Editorial and news coverage | High trust, and durable long after publication. | Slow to earn, but a single good piece keeps paying. |
Notice that four of the five are pages you do not own. That is the uncomfortable centre of this whole topic.
The three reasons brands score zero
A zero is almost always one of three things, and they need completely different responses.
1. Nobody outside your own domain describes you. You have a website, maybe a blog, and essentially no third-party footprint. The model has nothing to corroborate, so it stays quiet. This is the most common cause in companies under about $5M ARR, and the fix is slow but well understood: get described, accurately, in places you do not own.
2. You are described, but in the wrong category. Sources exist, but they file you somewhere buyers are not looking. You call yourself a "revenue intelligence platform" and buyers ask for "sales forecasting software". The model is answering correctly. You are simply not in the category being asked about. This one is often fixable in weeks, because it is a language problem rather than an authority problem.
3. Your category is dominated by a handful of incumbents. Generated answers name three to six companies. In a mature category, those slots are taken by brands with fifteen years of reviews. This is real and you should not pretend otherwise. The response is not to fight for the head term but to win the qualified variants, the "for [industry]", "for [company size]", "that integrates with [platform]" questions where specificity beats scale.
Diagnosing which of the three applies is the entire value of the first check. Treating cause three as if it were cause one wastes a year.
How this differs from your Google rankings
A lot of confusion comes from assuming these two move together. They overlap, but not reliably.
| What differs | Google ranking | ChatGPT visibility |
|---|---|---|
| The unit | A position in an ordered list | Named or not named. No eleventh place. |
| Consistency | Broadly stable hour to hour | Varies between runs and between users |
| Main lever | Relevance and links to a page | Consistent third-party description of a brand |
| Who you compete with | Whoever ranks for the keyword | The three to six names the model considers safe |
| Measurement | Search Console, free and exact | No console exists. You have to sample. |
| Speed of change | Days to weeks | Weeks for facts, months for recommendation |
Ranking first on Google helps, particularly because it usually means the corroborating pages exist. It guarantees nothing. Plenty of companies rank well and are never named, and a few are named constantly while ranking poorly, because the conversation about them happens somewhere other than their own site.
Turning observations into a number you can report
Mention rate is the headline, but on its own it is thin. Three numbers from the same data set make a far better report.
- Mention rate. Observations naming you, divided by total observations.
- Average position. When named, where in the list. First and fifth are not the same outcome.
- Accuracy rate. Of the answers naming you, how many got the facts right. This is the one executives react to, because a wrong price is a lost deal rather than a missed impression.
Run the identical prompt set for your two or three main competitors while you are in there. It roughly doubles the work and multiplies the value, because "we are named in 20 percent, the market leader in 70 percent" is a business statement. "We are at 20 percent" is a curiosity.
A worked example, end to end
Abstract instructions are easy to nod along to and hard to act on, so here is the whole thing with numbers.
A Series A company sells scheduling software to healthcare clinics. They build ten prompts, none containing their name, and run each three times in temporary chats over one week. Thirty observations.
The raw result: named in 4 of 30, so a mention rate of 13 percent. Of those four, they are listed fourth, fifth, fifth and third, so an average position of 4.25. In two of the four answers the model states they do not offer an API. They shipped one eighteen months ago, so the accuracy rate is 50 percent.
They run the same ten prompts for their two main competitors. One is named in 22 of 30, the other in 9 of 30.
Now the numbers say something. Thirteen percent against seventy-three percent is not a content gap, it is an authority gap. Being listed fourth or fifth means when they do appear, they appear as an afterthought. And half the time they are named, buyers are told something false that would disqualify them from any deal with an integration requirement.
The prioritisation falls out of the data. The API error is the emergency, because it is actively costing deals, it is factual rather than subjective, and it is fixable this week by correcting their own product page and the two review profiles the model kept citing. The 13 percent is the quarter-long project. The position problem solves itself if the other two do.
Notice what the exercise did not require: a tool, a budget, or permission. It required ninety minutes and a spreadsheet, and it produced a ranked action list with a business case attached to each item.
What good looks like
There is no universal benchmark, and anyone quoting one is guessing, because the number depends entirely on how crowded your category is. As a rough orientation for a B2B software category:
- 0 percent. You are invisible. Diagnose which of the three causes above applies before spending anything.
- 1 to 15 percent. You exist but you are the afterthought, named only on the most specific prompts.
- 15 to 40 percent. Genuine presence. Usually means real third-party coverage exists and the work now is position and accuracy.
- 40 percent and above. You are a default answer in your category. Protect it and watch accuracy.
Treat those bands as orientation, not as a target. The only comparison that means anything is your own number over time, measured the same way each time, and your competitors' numbers on the identical prompts.
ChatGPT visibility is not the whole picture
ChatGPT has the most users, so it is the sensible place to start. It is not the only place your buyers are.
Perplexity shows its sources inline, which makes it far more useful for diagnosis even though fewer people use it. Google AI Overviews reach anyone who searches at all, whether or not they ever open a chatbot. Gemini often disagrees with AI Overviews despite sharing an index. Microsoft Copilot sits inside the software your enterprise buyer has open right now, and is the engine most often left out of monitoring entirely.
Being strong in ChatGPT and absent everywhere else is a real and fairly common failure mode. Once you have your ChatGPT baseline, widen it. Our guide to AI brand monitoring covers running this across engines properly, including verified pricing for every major tool.
Tools that check ChatGPT visibility
You do not need one to start. You may want one once you know what you are tracking. Entry pricing verified from vendor pages in September 2026:
- Free checkers, including ours, for a first read across engines with no commitment
- AthenaHQ Essential, free tier with $25 of credit across five engines
- Otterly.ai, from $29 a month for 15 tracked prompts across four engines
- HubSpot AEO, $50 a month, capped at 25 prompts and without Google coverage, reviewed in full in our 28-day HubSpot AEO review
- Peec AI, from $95 a month for 50 prompts across three chosen models
- Scrunch, from $250 a month for 125 prompts, with a 7-day trial
- AthenaHQ Starter, $295 a month for 3,600 credits across 11 models
One warning on credit-based pricing. A credit is normally one AI response, so the cost scales with prompts multiplied by models multiplied by how often each is refreshed. Breadth of model coverage consumes quota fast, which means the plan that looks widest can track the fewest questions. Decide your prompt set before choosing a plan. Fuller comparisons sit in our roundups of the best ChatGPT SEO tools and the best AI visibility tools.
What to do with a bad score
Resist the urge to publish thirty blog posts. It is the standard response and it is mostly wasted, because your own content is the weakest of the signals involved.
First 30 days: fix what you control. Make sure your site states plainly what you do, for whom, in the category language buyers actually use. Un-gate pricing if you can. Correct anything factually wrong that a model might be reading. Publish llms.txt. None of this will move the number by itself, but everything after it depends on it being right.
Days 30 to 90: get described elsewhere. Complete and refresh review profiles. Make sure you appear, accurately, on the comparison pages that already rank for your category, including ones you do not own. Participate honestly where your buyers talk. This is where the number actually starts moving.
Beyond 90 days: build the thing nobody can copy. Original research, real data, a point of view that gets cited because it is worth citing. This is what turns you from a name that appears into a name that gets recommended first.
Re-run the identical check monthly. Same prompts, same temporary-chat method, same spreadsheet. The trend is the only thing that matters, and changing the method destroys the trend.
If you would rather not run this yourself, this is exactly what we do at Arobis AI. Our approach to AI search demand generation for B2B SaaS is built around moving the number rather than watching it, and our plans and pricing are public.
Mistakes that waste a quarter
- Checking logged in. The big one. Everything downstream is corrupted by it.
- Checking once. One answer is a sample of one from a distribution that varies.
- Asking "what is [your company]". That tests brand knowledge, not recommendation, and it will always look better than reality.
- Changing the prompts each month. Destroys comparability, which was the entire point.
- Publishing volume instead of earning corroboration. Your blog is the weakest signal in the system.
- Ignoring accuracy. Being named alongside a wrong price is worse than not being named.
- Expecting weeks. Facts move in weeks. Recommendation moves in months.
The vocabulary, defined
This category invents terms faster than it agrees on them, and vendors use the same word to mean different things. These are the definitions used throughout this guide.
Mention rate. The share of observations in which your brand is named at all. Observations, not prompts: ten prompts run three times is thirty observations.
Position. Where you appear in the list of companies within a single answer. Being named third of six is materially different from being named first, because readers stop early.
Share of voice. Your mentions as a proportion of all brand mentions across the same prompt set. This is the number that makes sense to executives, because it is competitive rather than absolute.
Citation. A source the engine references to support its answer. ChatGPT cites sparingly, Perplexity cites almost everything, which is why Perplexity is more useful for diagnosis.
Hallucination. A confident statement about you that is false. Treat it as a defect with a deadline, not as a curiosity.
Sentiment. Whether the description of you is positive, neutral or negative. Usually less actionable than accuracy, because neutral-but-correct is a perfectly good outcome.
Prompt set. The fixed list of questions you test. Fixed is the operative word. Changing it invalidates every comparison you have made.
Clean session. A temporary chat or logged-out window with no memory or history. The only condition under which a check is worth recording.
Frequently asked questions
What is ChatGPT visibility?
ChatGPT visibility is how often ChatGPT names your brand when someone asks a question your product could answer, and how accurately it describes you when it does. It is not a ranking position, because generated answers name a handful of companies rather than listing results in order.
How do I check my ChatGPT visibility for free?
Write eight to ten prompts a buyer who has never heard of you would ask, open a temporary chat so history and memory are excluded, run each prompt, and record whether you were named, in what position, whether anything was wrong, and what was cited. Repeat twice more on different days and divide mentions by total observations.
Why does ChatGPT recommend my company to me but not to my customers?
Because you are almost certainly signed in to an account with months of history about your own company. Memory and personalisation make the model far more likely to name you specifically to you. Always check in a temporary chat or logged out.
Does ChatGPT visibility depend on my Google rankings?
Not directly. Strong rankings often correlate with visibility because they usually mean third-party coverage exists, but ChatGPT weights consistent description of your brand across many independent sources rather than the relevance of a single page. Companies rank first and go unmentioned regularly.
How long does it take to improve ChatGPT visibility?
Factual corrections on pages models already read can show up within weeks. Changing whether you are recommended at all depends on third-party coverage and typically takes three to six months. Anyone promising faster is selling something.
Can I pay ChatGPT to recommend my brand?
No. There is no paid placement inside organic ChatGPT answers. Advertising products exist around AI search, but the recommendation itself is not purchasable, which is precisely why earned coverage matters so much.
How many prompts should I track?
Eight to ten is enough for a first read. Twenty or more gives a stable monthly number. What matters more than count is that they are genuinely different buying questions rather than rewordings of one, and that they stay fixed so the trend means something.
Does ChatGPT browse the web for these answers?
Sometimes. It may answer from training alone or retrieve live pages depending on the question and settings. This is one reason the same prompt returns different answers at different times, and another reason single checks are unreliable.
Should I use my brand name in the test prompts?
Only when you are deliberately testing brand knowledge. For recommendation testing, which is what matters commercially, never include your name. You are trying to find out whether the model brings you up unprompted.
What is a good ChatGPT mention rate?
It depends entirely on category competitiveness, so there is no universal figure. As orientation in B2B software, anything above roughly 40 percent means you are a default answer, and 15 to 40 percent means real presence. Your own trend and your competitors' numbers on identical prompts matter far more than any benchmark.
Is ChatGPT visibility the same as AI visibility?
No. ChatGPT is one engine. AI visibility covers ChatGPT, Perplexity, Google AI Overviews, Gemini, Claude and Copilot, which frequently disagree with each other. Being visible in one and absent in the rest is common.
Do I need a paid tool to track this?
Not to start. A spreadsheet and a fixed prompt set will give you a defensible baseline. Paid tools buy scheduling, scale and history rather than access to anything you could not check yourself.
Does ChatGPT visibility differ by country?
Yes. The same prompt asked from different countries can return different companies, because regional sources and language differ. If you sell internationally, test from each major market rather than assuming one reading covers all of them.
What if a competitor is named and we are not, but we are the better product?
Being better is not a signal a model can read. It reads how consistently and credibly you are described by others. A worse product with ten years of reviews will beat a better product with none, which is unfair and also fixable.
How is this different from checking Google AI Overviews?
AI Overviews sit on top of Google's index and are closer to the SEO you already do, so they often move faster if you already rank. ChatGPT draws more broadly on how the web describes you overall. The two frequently disagree, so track them separately.
The short version
ChatGPT visibility is binary in a way search never was. You are in the answer or you are not, and the answer names very few companies.
Check it properly, which means in a temporary chat, with prompts a stranger would type, repeated enough times to be a measurement rather than an anecdote. Most brands who do this for the first time get a worse number than they expected, and that is useful information rather than bad news.
Then work on the thing that actually drives it, which is being described consistently and accurately by people who are not you. That is slower than publishing another blog post, and it is the only thing that reliably works.



