The short answer
AI assistants find products through search crawlers that respect robots.txt, through pages users ask them to open, and through ordinary search indexes. Let the search crawlers in, put your key facts in HTML that loads without JavaScript, and make pricing and docs easy to fetch. In our check of 182 AI product sites, only 4 blocked an AI search crawler, but 13% served homepages that are nearly empty without JavaScript and another 13% turned away a plain automated request.
| Crawler | Company | Used for | Follows robots.txt | Sites blocking it (of 168) |
|---|---|---|---|---|
| GPTBot | OpenAI | Training data | Yes | 9 |
| OAI-SearchBot | OpenAI | ChatGPT search results | Yes | 3 |
| ChatGPT-User | OpenAI | Pages a user asks ChatGPT to open | May not apply | Not checked |
| ClaudeBot | Anthropic | Training data | Yes | 7 |
| Claude-SearchBot | Anthropic | Search quality in Claude | Yes | 2 |
| PerplexityBot | Perplexity | Perplexity search results | Yes | 3 |
| Perplexity-User | Perplexity | Pages fetched to answer a question | Generally not | Not checked |
| Google-Extended | Control token for Gemini training and grounding, not a separate crawler | Yes | 6 |
How AI assistants find products
AI companies run separate crawlers for separate jobs. OpenAI's GPTBot collects training data, while OAI-SearchBot fetches pages for ChatGPT's search results; Anthropic and Perplexity split their crawlers in the same way. Blocking a training crawler does not remove you from AI search, but blocking a search crawler does.
A third kind of request comes from users. When someone asks ChatGPT or Perplexity to open a page, the assistant fetches it on their behalf, and the companies say robots.txt may not apply to those requests.
Google's AI Overviews and AI Mode use the regular Search index. Google says there are no extra requirements and no special files or markup needed: a page that is indexed and eligible for a snippet can be used as a supporting link.
What we checked
On October 2, 2026 we checked the websites of 182 AI products in the SOTA catalog, counting each site once and leaving out products that live on a section of another company's site. For each site we read robots.txt and tested whether common AI crawlers may fetch the homepage, looked for /llms.txt and a sitemap, and fetched the homepage to record its title, description, share image, structured data and how much text it contains before JavaScript runs.
A robots.txt rule is a request, and bot-protection services can still challenge a crawler that robots.txt allows. Our homepage requests were plain HTTP requests, so a site that refused them may still let verified search crawlers through. The full results are available as a CSV.
Few AI products block AI search
Of the 168 sites with a readable robots.txt, 12 (7%) block at least one AI crawler, and most of them block training rather than search. Seven sites, including Canva, Lovart, Meshy and Muse, block OpenAI's training crawler while still allowing its search crawler: they opt out of training without leaving ChatGPT search.
Only 4 sites block an AI search crawler, and all four are AI assistants themselves: ChatGPT's site blocks PerplexityBot, Claude's blocks OpenAI's search crawler, and Meta AI and Poe block several. For a product that wants to be recommended, blocking AI search crawlers works against you.
The bigger risk: pages crawlers cannot read
20 of the 158 homepages we could load (13%) contain fewer than 500 characters of text before JavaScript runs, so a crawler that does not render JavaScript sees an almost empty page. Google renders JavaScript, but its own documentation notes that not all bots can, and recommends server-side rendering or prerendering.
Another 24 sites (13%) refused a plain automated request with an error or a challenge, whatever user agent we sent. Bot protection usually lets verified search crawlers through, but it is worth confirming that the AI crawlers you care about are allowed, and that your pricing and docs pages are not behind the same wall.
llms.txt is common, but optional
101 of the 182 sites (55%) publish an /llms.txt file, a Markdown summary of the site with links to its key pages. It was proposed in 2024 and is not an official standard, and Google says no such file is needed for its AI features.
It is cheap to add and useful for agents that read documentation, so it is worth publishing once your main pages are in order. Treat it as a guide to pages that already work, not a replacement for them.
Structured data describes the company more often than the product
98 of the 158 homepages we loaded (62%) include JSON-LD structured data. Most of it describes the organization and the website; only 48 (30%) describe the product itself as a SoftwareApplication, WebApplication or Product, with details such as category, platform and price.
Product-level structured data gives search engines and assistants unambiguous facts to quote, such as what the product is, which platforms it runs on and what it costs.
A short checklist
Allow OAI-SearchBot, Claude-SearchBot, PerplexityBot and Googlebot in robots.txt. If you do not want your content used for training, block the training crawlers separately.
Make sure your homepage, pricing page and docs show their key facts in the HTML that loads first, and that bot protection lets verified crawlers through.
State the facts an assistant needs in plain sentences: what the product does, who it is for, what it costs and its limits. Keep your product name and URLs stable, and redirect old ones after a rename.
Add product-level structured data and, optionally, an llms.txt that points to your best pages.
Sources & method
Official documentation was checked on 2026-10-02. Selection advice is our assessment. Interfaces and requirements may change; check the linked source for the version you intend to use.
- Dataset (CSV): crawler access and metadata for 182 AI product sites, checked October 2, 2026
- OpenAI: Overview of OpenAI crawlers
- Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?
- Perplexity crawlers
- Google's common crawlers (Google-Extended)
- Google: AI features and your website
- The /llms.txt proposal
- Google: JavaScript SEO basics
How SOTA reviews projects · Suggest a correction · Download the evaluation checklist