Every scan run through the free Arobis AI Visibility Checker tests whether ChatGPT, Gemini, Claude and Perplexity can actually read a website, and how well it is set up to be cited. This report aggregates all 1,188 scans of 1,169 websites from 1 June to 18 September 2026, with the full dataset free to download and cite.
The Website AI Readiness Report is a free, monthly benchmark of how ready real websites are for AI search. It is built by Arobis AI, an AI Search Demand Generation agency for B2B SaaS, from the scans people run through our free AI Visibility Checker. No survey, no estimates: every number comes from a live, automated scan of a real website.
1,188 scans of 1,169 websites, aggregated into one dataset. Each scan tests whether the crawlers behind ChatGPT, Gemini, Claude and Perplexity can read the site, and how well its structured data, headings, content and llms.txt set it up to be cited.
Buyers now ask AI assistants for shortlists, and the engines answer from what they can crawl and understand. This report shows where websites actually fail that test today, which signals separate the sites that pass, and how fast adoption is moving month by month.
Compare your own checker score with the tables below. If you write about AI search, quote any statistic with a link back: the data is CC BY 4.0 and every finding is written as a single citable sentence. If you want the gaps closed for you, that is the work Arobis AI does for B2B SaaS.
Jump to a section
Each statistic below is a single, citable sentence. All figures come from the 1,188 scans in this edition unless a source is named.
Who is in this sample. These are not the top 1,000 websites. They are 1,169 websites whose owners or marketers chose to check their AI visibility, which makes this the most AI-aware slice of the web we know of. Treat the adoption numbers as a ceiling for the wider web, not an average. See the methodology.
The AI Readiness Score is a 0 to 100 composite of six things an AI engine needs before it can read, understand and cite a website: crawler access, structured data, semantic structure, content and entity signals, llms.txt, and technical health. The full weighting is in the methodology. Across 1,188 scans the mean is 77.2 and the median is 83.
| Score band | Label | Websites | Share |
|---|---|---|---|
| 0 to 39 | Needs work | 73 | 6.1% |
| 40 to 69 | Getting there | 283 | 23.8% |
| 70 to 89 | Strong (AI-ready) | 411 | 34.6% |
| 90 to 100 | Excellent | 421 | 35.4% |
70.0% of websites clear the 70-point AI-ready line, which is the threshold for the green Verified AI-Ready badge. The distribution is top-heavy because the sample is self-selected: people who run an AI visibility check tend to have already invested in their site. Even so, 30.0% of these AI-aware websites still fall short, and 6.1% score below 40.
| Score | Websites | Share |
|---|---|---|
| 0 to 9 | 5 | 0.4% |
| 10 to 19 | 20 | 1.7% |
| 20 to 29 | 21 | 1.8% |
| 30 to 39 | 27 | 2.3% |
| 40 to 49 | 58 | 4.9% |
| 50 to 59 | 102 | 8.6% |
| 60 to 69 | 123 | 10.4% |
| 70 to 79 | 154 | 13.0% |
| 80 to 89 | 257 | 21.6% |
| 90 to 99 | 301 | 25.3% |
| 100 | 120 | 10.1% |
Most AI crawler studies read robots.txt and stop there. We do that too, for seven crawler tokens, but we also fetch every homepage live with each engine's real crawler user agent and check whether a usable page comes back. The two measurements disagree by a wide margin.
89.8% of websites publish a robots.txt file. Explicit AI crawler blocking is rare in this sample: GPTBot is disallowed on 4.8% of sites, and only 5.1% of sites disallow any of the seven AI tokens we check. That is far below the 19 to 25% GPTBot block rates reported for the top 1,000 to 5,000 websites by Presenc AI and HasData, which fits: this sample is made of businesses that want to be found by AI, not publishers protecting content.
| Crawler token | Operator | Sites blocking it | Share of 1,188 |
|---|---|---|---|
| GPTBot | OpenAI (training and browsing) | 57 | 4.8% |
| OAI-SearchBot | OpenAI (ChatGPT search) | 14 | 1.2% |
| ChatGPT-User | OpenAI (user-triggered fetches) | 16 | 1.3% |
| ClaudeBot | Anthropic (Claude) | 51 | 4.3% |
| Claude-Web | Anthropic (legacy token) | 17 | 1.4% |
| Google-Extended | Google (Gemini training and grounding) | 53 | 4.5% |
| PerplexityBot | Perplexity | 18 | 1.5% |
Blocked means the robots.txt group for that token disallows the site root. 10 sites (0.8%) could not be evaluated because robots.txt was unreachable.
For the 1,108 scans with live probe data, we requested the homepage over HTTPS with each engine's crawler user agent. Full means a 200 response with a real page body. HTTP only means HTTPS was refused but plain HTTP served the page. None means the crawler got a block page, an error, an empty body, or nothing within seven seconds.
| AI engine | Probed as | Full page | HTTP only | No usable page |
|---|---|---|---|---|
| Gemini | Googlebot | 92.8% | 0.5% | 6.7% |
| Perplexity | PerplexityBot | 89.6% | 0.5% | 9.8% |
| Claude | ClaudeBot | 81.0% | 0.6% | 18.2% |
| ChatGPT | GPTBot | 79.7% | 0.5% | 19.8% |
ChatGPT's crawler is turned away by 19.8% of websites and Claude's by 18.2%, while Googlebot, which feeds Gemini and AI Overviews, gets a full page from 92.8%. The gap is not a robots.txt decision. It is bot-protection defaults in CDNs and firewalls that ship with "AI scraper" rules switched on, which treat GPTBot and ClaudeBot as threats and Googlebot as a guest.
| AI engine | Sites whose robots.txt allows it | Blocked it on a live fetch anyway | Silent block rate |
|---|---|---|---|
| ChatGPT | 1,045 | 191 | 18.3% |
| Claude | 1,050 | 174 | 16.6% |
| Perplexity | 1,084 | 100 | 9.2% |
| Gemini | 1,050 | 67 | 6.4% |
18.3% of websites that explicitly permit OpenAI's crawlers in robots.txt still return nothing usable to GPTBot. Their owners believe they are open to ChatGPT. They are not. This is the single most common AI visibility problem we find, and the owner almost never knows about it because every human browser test passes.
Pattern worth naming: 4.9% of websites serve full pages to Googlebot and PerplexityBot but block both GPTBot and ClaudeBot. That exact signature is the fingerprint of a managed bot-protection rule set, not a deliberate policy. Sites where GPTBot got no page average 57.6 points; sites that served it a full page average 82.7.
llms.txt is a plain-text file at the site root that hands AI systems a curated map of what a site is about. No major engine has formally committed to reading it, yet adoption among sites that care about AI visibility is climbing fast. 45.6% of the websites in this report publish one, and the rate among newly scanned sites has risen every month.
| Month scanned | Websites scanned | With llms.txt |
|---|---|---|
| June 2026 | 253 | 37.5% |
| July 2026 | 313 | 40.6% |
| August 2026 | 130 | 49.2% |
| September 2026 (to 18 Sep) | 492 | 52.0% |
| Sample | llms.txt adoption | Basis |
|---|---|---|
| Websites that ran an AI visibility check (this report) | 45.6% | 1,169 websites, Jun to Sep 2026 |
| SE Ranking, 300,000 domains | 10.1% | 2026, reported by SE Ranking |
| Rankability, top 1,000 websites | 8.7% | June 2026 |
| Rankability, top 10,000 websites | 5.6% | June 2026 |
The gap is the point. In the general web, llms.txt is a single-digit phenomenon. Among businesses actively working on AI visibility it is approaching a coin flip, and platform defaults are part of the story: several website builders and e-commerce platforms now generate the file automatically. If you want a free one, use the llms.txt generator.
Does it correlate with readiness? Sites with llms.txt average 90.0 points; sites without average 66.5. Be careful with that number: llms.txt is worth exactly 10 points in our model, so 10 of the 23.5-point gap is mechanical. The other 13.5 points are real: teams that bother with llms.txt have usually also fixed their schema, headings and crawler access.
JSON-LD is how a website tells machines what it is in unambiguous terms. 75.3% of homepages include at least one JSON-LD block, which sounds healthy until you look at what is inside. Only 62.3% declare an Organization (or a subtype such as LocalBusiness or ProfessionalService), which is the one entity that lets an AI engine connect a page to a named company. FAQPage, the type most closely tied to answer-style citations, appears on 15.8%.
| Schema type | Websites | Share of 1,188 |
|---|---|---|
| WebSite | 727 | 61.2% |
| Organization | 675 | 56.8% |
| WebPage | 479 | 40.3% |
| BreadcrumbList | 298 | 25.1% |
| ImageObject | 268 | 22.6% |
| FAQPage | 188 | 15.8% |
| Person | 171 | 14.4% |
| Article | 105 | 8.8% |
| LocalBusiness | 104 | 8.8% |
| SoftwareApplication | 73 | 6.1% |
| Place | 57 | 4.8% |
| ItemList | 47 | 4.0% |
| Service | 47 | 4.0% |
| ProfessionalService | 36 | 3.0% |
| VideoObject | 27 | 2.3% |
A homepage can declare several types, so shares do not sum to 100%. WebSite and WebPage are frequently emitted by CMS plugins without any human input, which is why they lead the table.
| Signal | Adoption | Average score with | Average score without | Gap |
|---|---|---|---|---|
| Any JSON-LD block | 75.3% | 86.4 | 49.2 | +37.2 |
| Organization schema | 62.3% | 88.5 | 58.6 | +29.9 |
| FAQPage schema | 15.8% | 92.5 | 74.3 | +18.2 |
Structured data carries 20% of the composite score, so part of each gap is built in. The size of the gaps is still telling: the 24.7% of websites with no JSON-LD at all average 49.2 points, deep in the "needs work" band, and they rarely fail on schema alone.
AI engines extract answers from readable text and clean heading structure. The table ranks ten homepage signals by how common they are. HTTPS is effectively universal. Everything an engine uses to understand who a site is and what it says is far patchier.
| Homepage signal | Websites with it | Score dimension |
|---|---|---|
| HTTPS | 99.1% | Technical |
| Meta description present | 89.4% | Semantic structure |
| robots.txt present | 89.8% | Crawler access |
| Canonical tag | 84.2% | Technical |
| About or company signal | 82.8% | Content and entity |
| Exactly one H1 | 63.6% | Semantic structure |
| JSON-LD present | 75.3% | Structured data |
| Open Graph image | 67.8% | Content and entity |
| Organization schema | 62.3% | Structured data |
| llms.txt | 45.6% | llms.txt |
| H1 count | Websites | Share | Average score |
|---|---|---|---|
| No H1 | 257 | 21.6% | 57.6 |
| Exactly one H1 | 756 | 63.6% | 84.3 |
| Two or more H1s | 175 | 14.7% | 75.3 |
21.6% of homepages have no H1, usually because the headline is an image, a styled div, or rendered by JavaScript after load. An AI crawler that does not execute scripts sees a page with no title statement at all.
| Words of readable text | Websites | Share | Average score |
|---|---|---|---|
| Under 100 words | 136 | 11.4% | 42.4 |
| 100 to 299 | 73 | 6.1% | 67.7 |
| 300 to 799 | 320 | 26.9% | 76.4 |
| 800 or more | 659 | 55.5% | 85.8 |
The median homepage carries 900 words. The 11.4% with fewer than 100 words are mostly JavaScript-rendered single-page apps and image-led landing pages: they look fine to a human and are close to empty for a crawler. The median server response time in the sample is 643 ms, so speed is rarely the problem. Text that exists in the HTML is.
The same signals matter differently to each engine, so every scan also produces four engine-specific scores. The ranking is stable across the whole sample: Gemini first, Perplexity a close second, then Claude and ChatGPT nearly nine points behind.
| AI engine | Average readiness score | Sites serving its crawler a full page | Heaviest-weighted signals |
|---|---|---|---|
| Gemini | 82.7 | 92.8% | Crawler access 28%, structured data 26%, technical 18% |
| Perplexity | 82.2 | 89.6% | Crawler access 28%, content and entity 26%, semantic structure 20% |
| Claude | 75.1 | 81.0% | Crawler access 28%, semantic structure 26%, content and entity 22% |
| ChatGPT | 74.0 | 79.7% | Crawler access 28%, structured data 24%, semantic structure 22% |
The spread is almost entirely a crawler access story. Gemini leans on Googlebot, which nearly every site welcomes. ChatGPT and Claude depend on GPTBot and ClaudeBot, which 19.8% and 18.2% of websites turn away. For a B2B company, that means the engine buyers use most for vendor research is the one most likely to be unable to read the vendor's own site.
Each row is the cohort of websites scanned in that month, so this is a trend in who shows up and what they have already done, not a re-scan of the same sites. Two things move in the right direction: llms.txt adoption climbs 14.5 points and explicit GPTBot blocking in robots.txt falls from 8.3% to 3.3%. One thing does not: the share of sites that turn GPTBot away on a live fetch is as high in September (21.9%) as it was in June (22.1%). Owners are opening the door in robots.txt while their infrastructure keeps it shut.
| Cohort | Scans | Avg score | llms.txt | JSON-LD | Organization schema | GPTBot disallowed in robots.txt | GPTBot got no page live |
|---|---|---|---|---|---|---|---|
| June 2026 | 253 | 71.5 | 37.5% | 64.4% | 51.4% | 8.3% | 22.1% |
| July 2026 | 313 | 77.2 | 40.6% | 76.0% | 66.5% | 6.4% | 16.6% |
| August 2026 | 130 | 79.5 | 49.2% | 83.1% | 63.8% | 7.7% | 16.2% |
| September 2026 (to 18 Sep) | 492 | 79.5 | 52.0% | 78.5% | 64.8% | 3.3% | 21.9% |
Live probe data starts on 1 June 2026 and covers 1,108 of the 1,188 scans; the last column uses that subset (172 scans in June).
No single fix makes a site AI-ready, but the signals stack. Sites that have done both of the two most visible pieces of work, publishing llms.txt and shipping JSON-LD, average 93.2 points. Sites with neither average 47.2, a 46.0-point gap that is far larger than the 30 points those two signals control directly.
| Has the site published | Average score | Typical band |
|---|---|---|
| llms.txt and JSON-LD | 93.2 | Excellent |
| llms.txt only or JSON-LD only | 90.0 / 86.4 | Strong |
| Neither | 47.2 | Getting there (242 sites, 20.4%) |
Read in order of impact, the data says: first make sure GPTBot and ClaudeBot get a real page (check your CDN and firewall, not just robots.txt); second, declare who you are with Organization schema and one clear H1; third, put real text in the HTML; and only then add llms.txt and FAQPage on top. That is also the order our free checker lists fixes in, and it is the first month of every Arobis AI engagement.
Source. Every figure comes from scans run through the free Arobis AI Visibility Checker between 1 June and 18 September 2026: 1,188 completed scans of 1,169 distinct websites. A handful of sites were scanned more than once; repeat scans are counted as separate observations. Scans that failed to reach the site are excluded.
What a scan does. The checker fetches the homepage, robots.txt and /llms.txt, parses the HTML for headings, metadata, Open Graph tags, JSON-LD and readable word count, and reads robots.txt for seven AI crawler tokens: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-Web, Google-Extended and PerplexityBot. It then fetches the homepage live, once per engine, using the real crawler user agent for GPTBot, Googlebot, ClaudeBot and PerplexityBot (Google-Extended is a robots.txt token only, so Gemini is probed with Googlebot). A probe counts as a full page when the server returns 200 with at least 5 KB of body over HTTPS; a challenge page large enough to pass that floor would be counted as full, so the silent-block figures are conservative.
The score. Six dimensions, each 0 to 1, weighted into a 0 to 100 composite: crawler access 25%, structured data 20%, semantic structure 20%, content and entity signals 15%, llms.txt 10%, technical health 10%. A robots.txt disallow is decisive for an engine; otherwise the live probe decides (full = 1, HTTP only = 0.5, none = 0). Bands: 0 to 39 needs work, 40 to 69 getting there, 70 to 89 strong, 90 to 100 excellent. Per-engine scores reuse the same dimensions with engine-specific weights, listed in the engine table above. The scoring is deterministic and identical for every site.
Bias, stated plainly. This is a self-selected sample of websites whose owners or marketers wanted to know their AI visibility. It over-represents marketing-aware small and mid-sized businesses, agencies and SaaS companies, and under-represents publishers, enterprises and the long tail. Adoption figures here are therefore an upper bound for the web at large. Correlations between a signal and the score are partly mechanical, because the signal contributes to the score; we say how much wherever it matters. 22.6% of sites were submitted with a www prefix.
| TLD | Websites | Share |
|---|---|---|
| .com | 629 | 52.9% |
| .ir | 59 | 5.0% |
| .in | 51 | 4.3% |
| .lt | 50 | 4.2% |
| .uk | 39 | 3.3% |
| .ai | 31 | 2.6% |
| .au | 29 | 2.4% |
| .app | 23 | 1.9% |
| .io | 20 | 1.7% |
| .ca | 16 | 1.3% |
| .co | 14 | 1.2% |
| .net | 14 | 1.2% |
Cadence. This page is refreshed on the first working week of every month with the new cohort added and the headline numbers, tables and dataset re-generated from the same queries. The edition and last-updated date appear at the top and in the page's structured data.
The aggregate dataset behind every table on this page is published as JSON under a Creative Commons Attribution 4.0 licence. Use it in articles, decks and research; link back to this page as the source. Individual scanned domains are not included.
Arobis AI (September 2026). Website AI Readiness Report 2026: What Real AI Visibility Scans Reveal. https://arobis.ai/ai-readiness-reportAccording to Arobis AI's analysis of 1,188 website scans, 19.8% of websites return no usable page to ChatGPT's crawler while only 4.8% block it in robots.txt.Journalists and researchers: for a custom cut of the data (by industry, platform or country) email ran@arobis.ai.
An AI readiness score measures how well a website is set up to be read, understood and cited by AI engines such as ChatGPT, Gemini, Claude and Perplexity. The Arobis AI version is a 0 to 100 composite of six dimensions: crawler access, structured data, semantic structure, content and entity signals, llms.txt and technical health. Across 1,188 scans in 2026 the average is 77.2.
It depends on how you measure. In robots.txt, only 4.8% of the websites in this report disallow GPTBot. On a live fetch with GPTBot's real user agent, 19.8% of websites return no usable page. Most ChatGPT blocking is done silently by CDN and firewall bot rules, not by robots.txt.
45.6% of the 1,169 websites in this report publish an llms.txt file, rising from 37.5% of sites scanned in June 2026 to 52.0% in September. Studies of the top 1,000 and top 10,000 websites report 8.7% and 5.6%, so adoption is far higher among businesses actively working on AI visibility than on the web at large.
Gemini. 92.8% of websites served a full page to Googlebot, the crawler that feeds Gemini and Google AI Overviews. Perplexity followed at 89.6%, Claude at 81.0% and ChatGPT at 79.7%.
No AI engine has formally committed to reading llms.txt, so treat it as low-cost insurance rather than a ranking factor. In this dataset sites with llms.txt score 90.0 on average against 66.5 without, but 10 of those points are the file's own weight in the model and the rest reflects that llms.txt adopters usually also have proper schema, headings and crawler access.
Every number comes from scans run through the free Arobis AI Visibility Checker between 1 June and 18 September 2026. The aggregate dataset is published as JSON under a Creative Commons Attribution 4.0 licence, so you can quote it, chart it and republish it with a link to this page. Individual domains are not shared.
Run the free Arobis AI Visibility Checker at arobis.ai/ai-visibility-checker. It performs the same scan used in this report, returns your score, the four engine scores and a prioritised fix list in about a minute, and sites scoring 70 or higher can embed a Verified AI-Ready badge.
The free checker runs the exact scan behind this report: live crawler probes for four AI engines, schema, headings, llms.txt and a prioritised fix list. Score 70 or higher and you can show the Verified AI-Ready badge on your site.
AI Search Demand Generation for B2B SaaS. Get recommended by ChatGPT, Gemini, Claude, Perplexity, Copilot and Google AI.