Submit projectSubmit

THE SOTA FIELD GUIDE

How to get your AI product found by ChatGPT, Claude and Perplexity

Which AI crawlers to allow, what they need to read on your site, and what 182 AI product sites actually do today.

By SOTA · Updated 2026-10-02 · Researched with AI assistance; facts checked against the linked sources · How we work

The short answer

AI assistants find products through search crawlers that respect robots.txt, through pages users ask them to open, and through ordinary search indexes. Let the search crawlers in, put your key facts in HTML that loads without JavaScript, and make pricing and docs easy to fetch. In our check of 182 AI product sites, only 4 blocked an AI search crawler, but 13% served homepages that are nearly empty without JavaScript and another 13% turned away a plain automated request.

Summary of the findings
CrawlerCompanyUsed forFollows robots.txtSites blocking it (of 168)
GPTBotOpenAITraining dataYes9
OAI-SearchBotOpenAIChatGPT search resultsYes3
ChatGPT-UserOpenAIPages a user asks ChatGPT to openMay not applyNot checked
ClaudeBotAnthropicTraining dataYes7
Claude-SearchBotAnthropicSearch quality in ClaudeYes2
PerplexityBotPerplexityPerplexity search resultsYes3
Perplexity-UserPerplexityPages fetched to answer a questionGenerally notNot checked
Google-ExtendedGoogleControl token for Gemini training and grounding, not a separate crawlerYes6

How AI assistants find products

AI companies run separate crawlers for separate jobs. OpenAI's GPTBot collects training data, while OAI-SearchBot fetches pages for ChatGPT's search results; Anthropic and Perplexity split their crawlers in the same way. Blocking a training crawler does not remove you from AI search, but blocking a search crawler does.

A third kind of request comes from users. When someone asks ChatGPT or Perplexity to open a page, the assistant fetches it on their behalf, and the companies say robots.txt may not apply to those requests.

Google's AI Overviews and AI Mode use the regular Search index. Google says there are no extra requirements and no special files or markup needed: a page that is indexed and eligible for a snippet can be used as a supporting link.

What we checked

On October 2, 2026 we checked the websites of 182 AI products in the SOTA catalog, counting each site once and leaving out products that live on a section of another company's site. For each site we read robots.txt and tested whether common AI crawlers may fetch the homepage, looked for /llms.txt and a sitemap, and fetched the homepage to record its title, description, share image, structured data and how much text it contains before JavaScript runs.

A robots.txt rule is a request, and bot-protection services can still challenge a crawler that robots.txt allows. Our homepage requests were plain HTTP requests, so a site that refused them may still let verified search crawlers through. The full results are available as a CSV.

The bigger risk: pages crawlers cannot read

20 of the 158 homepages we could load (13%) contain fewer than 500 characters of text before JavaScript runs, so a crawler that does not render JavaScript sees an almost empty page. Google renders JavaScript, but its own documentation notes that not all bots can, and recommends server-side rendering or prerendering.

Another 24 sites (13%) refused a plain automated request with an error or a challenge, whatever user agent we sent. Bot protection usually lets verified search crawlers through, but it is worth confirming that the AI crawlers you care about are allowed, and that your pricing and docs pages are not behind the same wall.

llms.txt is common, but optional

101 of the 182 sites (55%) publish an /llms.txt file, a Markdown summary of the site with links to its key pages. It was proposed in 2024 and is not an official standard, and Google says no such file is needed for its AI features.

It is cheap to add and useful for agents that read documentation, so it is worth publishing once your main pages are in order. Treat it as a guide to pages that already work, not a replacement for them.

Structured data describes the company more often than the product

98 of the 158 homepages we loaded (62%) include JSON-LD structured data. Most of it describes the organization and the website; only 48 (30%) describe the product itself as a SoftwareApplication, WebApplication or Product, with details such as category, platform and price.

Product-level structured data gives search engines and assistants unambiguous facts to quote, such as what the product is, which platforms it runs on and what it costs.

A short checklist

Allow OAI-SearchBot, Claude-SearchBot, PerplexityBot and Googlebot in robots.txt. If you do not want your content used for training, block the training crawlers separately.

Make sure your homepage, pricing page and docs show their key facts in the HTML that loads first, and that bot protection lets verified crawlers through.

State the facts an assistant needs in plain sentences: what the product does, who it is for, what it costs and its limits. Keep your product name and URLs stable, and redirect old ones after a rename.

Add product-level structured data and, optionally, an llms.txt that points to your best pages.

Sources & method

Official documentation was checked on 2026-10-02. Selection advice is our assessment. Interfaces and requirements may change; check the linked source for the version you intend to use.

How SOTA reviews projects · Suggest a correction · Download the evaluation checklist

Explore the projects