When a B2B buyer opens an AI assistant and types “what’s the best CRM for a mid-market sales team,” the answer that comes back is not a search page. It is a shortlist — three to seven named vendors, delivered as confident fact, before the buyer visits a single website. We wanted to know which vendors make that list, which ones are quietly left off, and how often the list is simply wrong. So we ran a controlled study across three core B2B SaaS categories — CRM, project management, and marketing automation — and measured Recommendation Frequency: the share of buying-intent prompts in which each vendor was actually recommended.
The results describe a market most SaaS teams cannot see. A small, stable set of incumbents is recommended in nearly every answer. Beloved, heavily funded challengers are recommended far less often than their brand awareness would predict. And in one category, a leading AI model recommended two products that no longer exist under those names — plus one product that never existed at all — with the same confidence it used to recommend HubSpot. This is the pre-website funnel in action: the buying decision narrows inside the AI answer, and most vendors never learn they were excluded.
Key findings: in the CRM category we tested, five vendors were recommended in 100% of buying-intent prompts while a functional competitor was recommended in just 6% — a 94-point recommendation gap inside the same question.
This is a single-model, directional first study, and we say so plainly below. But the pattern it exposes is the one that matters for every SaaS go-to-market team right now: AI visibility for SaaS is no longer about ranking on a results page. It is about whether an AI model names you when a buyer asks for a recommendation. For the wider context on how fast this shift is moving, see our State of AI Search Visibility research.
Finding 1: Recommendation is winner-take-all

The clearest signal in the data is concentration. When we asked a leading AI model to recommend CRM software, the same handful of names appeared again and again, and the rest of the market barely registered. Five vendors were recommended in every single prompt. Two more — Salesforce and Microsoft Dynamics 365 — appeared in most answers but not all. Below that tier, Recommendation Frequency collapsed toward zero.
Read that table the way a buyer experiences it. A prospect asking for a CRM recommendation will almost certainly hear about HubSpot, Zoho, Freshsales, Pipedrive, and Copper. They will usually hear about Salesforce. They will frequently hear about Dynamics. And they will almost never hear about Salesmate — a real, functional CRM — because at 6% Recommendation Frequency it is, for practical purposes, absent from the conversation. The buyer is not choosing to ignore it. The model never surfaced it.
What is striking is that Salesforce, the category’s revenue leader, did not top the list on Recommendation Frequency. At 94% it trailed five vendors on a pure recommendation basis — a reminder that market share and AI recommendation share are different currencies. We explored that specific dynamic in more depth in our HubSpot vs. Salesforce AI visibility study, and the CRM data here reinforces it: being the biggest does not guarantee being recommended most.
The mechanism behind winner-take-all is not mysterious. AI models assemble recommendations from patterns in the text they were trained on and the sources they retrieve — consistent, corroborated mentions across many contexts. Vendors that are described the same way across thousands of pages become the “safe” answer; vendors that lack that density fall off. We break down the specific inputs models weigh in our analysis of the five signals AI engines use to recommend brands. The practical consequence is brutal for challengers: the long tail is not slightly disadvantaged, it is invisible, and the vendors in it usually have no idea. If that describes your product, the first step is understanding why your brand doesn’t appear in AI answers. This concentration effect is not unique to CRM — we see the same handful-of-winners shape across categories in our review of the companies dominating AI discovery.
Finding 2: Brand love and funding did not buy recommendations

If Finding 1 is about incumbents winning, Finding 2 is about who loses — and the answer surprised us. The vendors most under-recommended relative to their market presence were not obscure. They were some of the most talked-about, best-funded, most-loved products in B2B software. In project management, the ranking by Recommendation Frequency looked almost inverted from what brand awareness would predict.
Consider what the market spends to build the brands at the bottom of this table. Monday.com is a public company that has poured enormous sums into advertising, yet it was recommended in 67% of prompts — behind Basecamp, a deliberately small, bootstrapped product that hit 100%. Notion, one of the most beloved tools of the last decade with a passionate user base and constant social buzz, landed at 39%. ClickUp, which built its growth on aggressive marketing and a “one app to replace them all” message, came in at just 11%. Brand love and ad budgets, it turns out, do not translate directly into AI recommendations.
The reason is that recommendation authority and marketing spend are not the same input. A model recommending project management software is not counting your billboards or your funding rounds; it is drawing on how consistently and authoritatively your product is described in the contexts it learned from. A tool can be adored by its users and still lack the corroborated, category-anchored footprint that makes a model reach for it by default. We unpack that distinction — and how models actually rank SaaS options — in our piece on how AI analyzes, recommends, and ranks SaaS products.
For marketing leaders, this is the uncomfortable part: the metrics that look healthy on a brand dashboard — awareness, sentiment, share of voice — can coexist with near-invisibility in AI answers. Recommendation Frequency is a separate scoreboard, and most teams are not yet tracking it. For a framework on bringing that number to the leadership table, our CMO guide to AI recommendation share lays out how to report it. The takeaway from the project management data is blunt: you can win the brand and still lose the shortlist.
“Every founder I talk to assumes that if buyers love the product and the brand is everywhere, the AI will recommend them. The data says otherwise. Visibility gets you seen. Recommendations get you chosen — and those are two different games. The vendors winning inside AI answers aren’t always the best-funded or best-loved; they’re the ones the model can describe with confidence, over and over. That’s a solvable problem, but only if you measure it.” — Ran Yosef, Founder, Arobis AI
Finding 3: AI recommends stale and fabricated products as fact

The third finding is the one that should give every SaaS operator — and every buyer — pause. In the marketing automation category, a leading AI model did not just misrank vendors. It recommended products that no longer exist under the names it used, and one product that never existed at all, presenting each with the same confident tone it used for HubSpot.
Autopilot was renamed Ortto in 2022. Sendinblue became Brevo in 2023. Both rebrands are years old, yet the model recommended the retired names as current options — sending a hypothetical buyer to look for products that, under those labels, no longer exist. Then there is “Autodesk Market Automation,” which is not a marketing automation product at all. It is an AI hallucination: a plausible-sounding vendor stitched together from fragments and recommended as if it were real. A buyer acting on that answer would go hunting for a product that has never existed.
Note also where Mailchimp landed. It is one of the most recognized brands in all of marketing, and at 72% Recommendation Frequency it trailed lesser-known incumbents like ActiveCampaign, Pardot, and Marketo, all at 100%. The same pattern from Finding 2 repeats: recognition is not recommendation. But the deeper issue in this category is trust. When a model confidently recommends a renamed product and a fabricated one inside the same answer, it reveals that AI recommendations are built from a snapshot of text that can be stale, incomplete, or wrong — and delivered without the hedging a human expert would use.
For vendors, the stale-data problem cuts both ways. If your product rebranded, changed positioning, or launched recently, the model may be recommending an outdated version of you — or not recommending you because it still “knows” the old you. For a fuller treatment of how models pick between brands and where these errors come from, see our breakdown of how the major AI engines actually choose which brands to mention. The correction is not to argue with the model; it is to build enough consistent, current, authoritative presence that the model’s next snapshot describes you accurately.
Methodology
We want to be precise about what this study is and is not, because credibility matters more than a big number. This is a single-model, directional first study, and every figure above comes from the same source.
A single model is a limitation, and we treat it as one. But Llama 3.3 is not a fringe system; it is one of the most widely deployed open models in production, which makes its recommendation behavior directly relevant to buyers using tools built on it. For the broader evidence base across the AI search shift, we maintain a running collection of 100+ AI search statistics, and this study is designed to be one honest data point within it. We would rather publish a transparent single-model study than an inflated cross-model claim we cannot stand behind.
What this means: the pre-website funnel
Put the three findings together and a new shape of the buyer journey comes into focus. The classic funnel assumed buyers discover vendors through search, content, and ads, then visit websites to evaluate. AI inserted a step before all of that. Now the first cut — who even makes the consideration set — happens inside the AI answer, before any website visit. We call this the Pre-Website Funnel, and it changes the job of every SaaS marketing team.
For vendors, the implication is direct: if you are not in the AI Shortlist, you are not in the deal. A buyer who hears five CRM names and yours is not among them will evaluate those five. Your beautifully optimized website, your demo, your pricing page — none of it gets a chance, because the buyer never arrives. Recommendation Frequency becomes an upstream revenue metric: it governs how many qualified buyers ever reach your funnel in the first place. This is why we frame the discipline as AI Search Demand Generation rather than a subset of SEO — the goal is not a ranking, it is a recommendation.
For buyers, the findings are a caution. An AI shortlist feels authoritative and complete, but this study shows it can omit strong options (the invisible long tail), under-weight excellent challengers (Notion, ClickUp), and include stale or fabricated vendors (Autopilot, Sendinblue, “Autodesk Market Automation”). Treat the AI answer as a confident intern’s first draft, not a verified market map. Verify that recommended products still exist under the names given, and ask specifically about categories or vendor types the model skipped.
The strategic response for vendors is not to game a model but to build durable recommendation authority — the consistent, corroborated, current presence that makes a model reach for you by default. That is a long game of category-anchored content, third-party corroboration, and accurate, everywhere-consistent descriptions of what you do. We lay out the playbook in our guide to building AI recommendation authority, and we structure client work around the same idea through The Arobis AI Search Demand Framework™. To see where the whole category is heading and how to prioritize, our State of AI Search Visibility report is the companion read to this study.
How to check your own brand
The worst position to be in is the one most vendors are in right now: invisible in AI answers and unaware of it. Because the buyer never arrives to tell you, the gap is silent. The fix starts with measurement — you cannot improve a Recommendation Frequency you have never looked at. Here is how to check your own brand in an afternoon. For the full step-by-step method, see our guide on how to check if AI recommends your brand.
Start with the extension if you want an immediate gut check — the Chrome extension is free and takes a minute to add. Move to the web-based AI visibility checker when you want to test specific buying prompts across your category. And when the stakes are high enough to warrant a full picture, our audit and plan options turn the findings into a prioritized roadmap. Whichever you choose, the goal is the same: know your number before your competitors define it for you.
Frequently asked questions
What is Recommendation Frequency?
Recommendation Frequency is the share of buying-intent prompts in which an AI model actually recommends a given vendor. In this study, each category used six standardized prompts run three times each (18 prompt-runs), so a vendor recommended in every run scores 100% and one recommended in a single run scores about 6%. It measures whether you get named when a buyer asks for a recommendation — a different and more decisive metric than traditional visibility.
Which AI model produced this data?
All of it came from one model: Meta’s Llama 3.3 70B, a leading, widely used AI model and the engine behind Arobis’s free tool. We did not test ChatGPT, Gemini, Claude, Perplexity, or Copilot, and we make no claims about their outputs. This is a transparent, single-model, directional first study. For the wider evidence base and where the category is heading, see our State of AI Search Visibility research.
Why were beloved tools like Notion and ClickUp recommended so rarely?
Because brand love and ad spend are not the inputs a model uses to recommend. Notion (39%) and ClickUp (11%) are widely used and heavily marketed, but recommendation authority comes from consistent, corroborated, category-anchored descriptions in the sources a model learned from — not from awareness or funding. A tool can be adored and still lack that footprint, which is why recognition and Recommendation Frequency diverge so sharply in the data above.
How can AI recommend a product that does not exist?
AI models generate text from patterns, not from a verified vendor database. When fragments combine plausibly, a model can produce a confident recommendation for something inaccurate — a renamed product (Autopilot, now Ortto; Sendinblue, now Brevo) or a fabricated one (“Autodesk Market Automation”). It is not lying; it is pattern-matching without a fact-check. That is why buyers should verify AI shortlists and vendors should keep their public descriptions accurate and current.
How do I find out if my brand is being recommended?
Start with the free Arobis Chrome extension to check as you browse, then run your category through the AI Visibility Checker to see who gets recommended for your buying prompts. For a complete competitive picture and a plan, review our audit options.
What should vendors do if they are invisible in AI answers?
Measure first, then build recommendation authority — the consistent, corroborated, current presence that makes a model reach for you by default. This is AI Search Demand Generation, and it is a distinct discipline from SEO because the goal is being recommended, not ranked. Our AI visibility for SaaS overview shows how it maps to revenue, and The Arobis AI Search Demand Framework™ turns the measurement into a prioritized plan.



