Guide
Training and search AI crawlers: which to allow in robots.txt
Published October 7, 2026
The companies behind AI assistants mostly run more than one crawler, each with its own name in robots.txt and its own job. One collects pages that may be used to train AI models. One gathers pages for the answers an AI search feature gives. One fetches a single page because a person asked an assistant about it. Blocking each costs a store something different, so the useful question is not whether to block AI crawlers but which ones, and robots.txt answers it one crawler at a time.
The kinds of AI crawler, and their robots.txt names
Each company documents its crawlers on its own site. What follows is what those pages said on October 7, 2026; each is linked at the foot of this page. The name in code is the robots.txt token, the word that goes after User-agent:, and crawlers match it without regard to case.
Training crawlers
These collect pages that may be used to train AI models. A Disallow for one asks its operator not to use your pages that way.
GPTBot(OpenAI). OpenAI says disallowing it "indicates a site's content should not be used in training generative AI foundation models".ClaudeBot(Anthropic). Anthropic says restricting it signals that a site's future materials should be excluded from its AI model training datasets.MistralAI-Training(Mistral AI), which Mistral says helps build datasets for training its generative AI models and is not used for search indexing or to answer live user questions.Meta-ExternalAgent(Meta), which Meta says crawls "for use cases such as training foundation AI models or improving products by indexing content directly": training, and more.Amazonbot(Amazon), which Amazon says is used to improve its products and services and "may be used to train Amazon AI models": training, and more.
Training controls that crawl nothing
Two tokens send no requests of their own. Their operators read them from robots.txt as instructions about pages their other crawlers fetch.
Google-Extended(Google) decides whether content Google crawls may be used to train future Gemini models, and to ground answers in Gemini Apps and on Vertex AI. Google says it does not affect a site's inclusion in Google Search and is not a ranking signal there.Applebot-Extended(Apple) decides whether pages Applebot crawled may train Apple's foundation models. Apple says pages that disallow it can still appear in its search results, and that its rules are not considered in ranking.
AI search crawlers
These read pages for the answers an AI search feature gives, and the links in them.
OAI-SearchBot(OpenAI), for ChatGPT's search features.Claude-SearchBot(Anthropic), which Anthropic says "navigates the web to improve search result quality for users".PerplexityBot(Perplexity), to surface and link websites in Perplexity's search results. Perplexity says it is not used to crawl content for AI foundation models.Meta-WebIndexer(Meta), to improve Meta AI's search results.Amzn-SearchBot(Amazon), for search in Amazon products such as Alexa. Amazon says it does not crawl content for generative AI model training.MistralAI-Index(Mistral AI), for Mistral's search, which answers questions in its Vibe assistant. Mistral says what it crawls is not used for training.DuckAssistBot(DuckDuckGo), which DuckDuckGo says crawls pages in real time for the AI-assisted answers in DuckDuckGo Search. DuckDuckGo says its data is not used to train AI models.
User-triggered fetchers
These do not crawl on their own. They fetch a page when a person asks an assistant a question that needs it, or hands it a link.
ChatGPT-User(OpenAI), for user actions in ChatGPT and custom GPTs.Claude-User(Anthropic), when people ask Claude questions.Perplexity-User(Perplexity), when people ask Perplexity a question.Meta-ExternalFetcher(Meta), which fetches links at a user's request, including to help AI complete tasks on websites for users.Amzn-User(Amazon), for example to fetch current information for an Alexa answer.MistralAI-User(Mistral AI), when people ask Vibe a question.- Google's user-triggered fetchers, among them Google-Agent (agents acting on a user's request) and the Gemini Notebook fetcher, have no robots.txt token: Google lists only the user-agent strings they send.
Search engines that also feed AI answers
Three search crawlers also feed AI answers: Google's AI Overviews and AI Mode, Copilot's answers from Bing's index, and the answers Apple's AI models give. No robots.txt token keeps a store out of those while leaving it in search. Google, Bing and Apple each offer controls set on the page instead, further down; all of Google's also change what its search results show, and so do Bing's nosnippet and data-nosnippet, and Apple's nosnippet.
Googlebot(Google). Google says its robots.txt rules for Googlebot are the control for Search, AI features included, and that a page needs only to be indexed and eligible to show in Search with a snippet to be eligible as a supporting link in AI Overviews or AI Mode.bingbot(Microsoft). Bing's guidelines say Bing and Copilot search experiences rely on the same crawling, indexing and ranking as Bing search, and name no separate crawler for Copilot.Applebot(Apple), for search in Spotlight, Siri and Safari. Apple says its data may also train Apple's foundation models (the control for that isApplebot-Extended) and add context to answers Apple's AI models generate. When robots.txt names Googlebot but not Applebot, Applebot follows Googlebot's rules.
What blocking each kind costs a store
Blocking GPTBot, ClaudeBot or MistralAI-Training asks its operator not to train on your pages, and each operator's search crawler is a separate token the block does not touch. OpenAI says its settings are independent of each other, and gives the example of a site that allows OAI-SearchBot so as to appear in search results while disallowing GPTBot. Blocking Applebot-Extended asks Apple not to use the pages Applebot reads to train its foundation models, and Apple says pages that disallow it can still appear in its search results.
Three tokens withdraw more than training. Google-Extended also covers grounding answers in Gemini Apps and on Vertex AI, though Google says it does not affect Google Search. Meta-ExternalAgent and Amazonbot also serve their operators' other products. Those two companies' search crawlers, Meta-WebIndexer and Amzn-SearchBot, are separate tokens, and Amazon says each of its user-agent settings is independent of the others.
Blocking an AI search crawler asks its operator to leave your pages out of that search feature. OpenAI says sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though they can still appear as navigational links. Anthropic says disabling Claude-SearchBot may reduce a site's visibility and accuracy in its users' search results. Perplexity recommends allowing PerplexityBot for a site to appear in its search results, and Amazon says permitting Amzn-SearchBot makes a site's content eligible for search experiences such as Alexa. DuckDuckGo says a block on DuckAssistBot opts a site out of being a source for its AI-assisted answers, and does not affect organic rankings or whether the site appears in its search results.
Blocking a user-triggered fetcher stops an assistant from reading your page when a shopper asks about it, wherever the operator honours the block (some say they may not; see below). Anthropic says disabling Claude-User prevents Claude from retrieving your content in response to a user's query, which may reduce your site's visibility for user-directed web search. OpenAI notes that ChatGPT-User plays no part in deciding what appears in ChatGPT search: that is OAI-SearchBot's job.
Blocking Googlebot or bingbot to keep pages out of AI answers shuts the store's pages off from those search engines too. Google says blocking Googlebot affects Google Search, Discover and every Search feature, and Google Images, Video and News; Bing's guidelines list blocking Bingbot among the things to avoid. Bing documents a tag that keeps a page's content out of Copilot answers while the page stays in its search results; Google's page controls limit what all of Google Search shows from a page, its AI features included (both below).
robots.txt for three common choices
One rule from the robots.txt standard, RFC 9309, decides how these are written. A crawler obeys the groups that name it, read together as one, and no others; the User-agent: * group is the fallback for crawlers no group names. Google puts it plainly: groups for a specific user agent and the * group are not combined. So a group that names a crawler replaces all of your * rules for that crawler, including the ones your platform wrote to keep crawlers out of carts, checkouts and admin pages.
Two operators document a fallback of their own. Apple says Applebot follows Googlebot's rules when a file names Googlebot but not Applebot, and Amazon says that when a file does not mention Amzn-SearchBot but allows other search bots, Amzn-SearchBot follows the rules given to them. A platform or CDN can also write named groups into your file: BigCommerce does for many of the crawlers here, and so does Cloudflare's managed robots.txt (both below).
To let a crawler in, leave it out of your file. A group such as User-agent: OAI-SearchBot with Allow: / takes that crawler off your * group, cart and checkout rules included. Give a crawler a group of its own only to block it.
Paste a snippet at the very top of the file, or below a group that has an Allow or Disallow rule of its own. RFC 9309 says a line it does not define, such as Sitemap:, must not end a group, and Google's parser reads only Allow:, Disallow: and User-agent: lines and ignores the rest. So a group just above the snippet with no rule of its own, such as a User-agent: line followed only by a Crawl-delay: or Sitemap: line, is read as part of the snippet's first group: that crawler would be kept out with it.
Before you paste, look for a group already in your file that names one of the snippet's crawlers and has an Allow line. RFC 9309 reads the groups that name a crawler together as one and uses the most specific matching rule, and where an Allow and a Disallow rule are equivalent, it says the Allow rule should be used. So an Allow: / there should win over the snippet's Disallow: / on every page, and a longer Allow path there, such as Allow: /products/, wins on the pages it matches. Edit that group instead of adding a second one for the same crawler: take out its Allow lines and give it Disallow: /. If that group also names a crawler the snippet does not name, take out only the User-agent: line of each crawler you are keeping out, so that the snippet's group is the only one naming it.
Choice 1: let every AI crawler in
Add nothing. A crawler no group names follows your * group, which already carries your platform's rules, and BigCommerce's help says no entries are needed to allow AI bots. Read your live file first (the last section says how): if a group in it disallows a crawler you now want in, delete that group, or, if a setting such as Cloudflare's managed robots.txt wrote it, turn that setting off instead, and read the file again afterwards (below). WordPress.com says that ticking its Prevent third-party sharing box adds known AI bots to the disallow list in a site's robots.txt (below); on a site it hosts, read the live file to see which bots it names. A CDN can also turn crawlers away whatever the file says (see the section on CDNs below).
Choice 2: keep training crawlers out
Add one group for each training crawler and training control above, each with Disallow: /, or edit a group your file already has for one of them, as above. Three of these withdraw more than training: Google-Extended also withdraws your pages from grounding in Gemini Apps and on Vertex AI, where Google gives the model content from its Search index as it answers, and Meta-ExternalAgent and Amazonbot serve their operators' other products too. Leave a group out if that other use matters more to you than the training it withdraws.
# AI training crawlers and training controls
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: MistralAI-Training
Disallow: /
User-agent: Meta-ExternalAgent
Disallow: /
User-agent: Amazonbot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /The snippet leaves every crawler it does not name, the AI search crawlers, the user-triggered fetchers and the search engines included, reading your file exactly as it did before, whichever group it followed: your * group, one your platform or CDN wrote for it, or the fallback its operator documents. Two platforms can change the rest of the file as you add it, so check it afterwards (below): on Shopify, compare it with the copy you saved, because Shopify says a template's output doesn't always match the file it generates, and on BigCommerce, check how your groups were merged.
Choice 3: keep the AI crawlers named here out, let search engines in
Add a group for each training crawler, training control, AI search crawler and user-triggered fetcher above, each with Disallow: /, or edit a group your file already has for one of them, as above. Every crawler the snippet does not name, Googlebot, bingbot and Applebot among them, reads your file exactly as before; on Shopify and BigCommerce, check the file afterwards, as for choice 2 (below).
# AI training crawlers and training controls
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: MistralAI-Training
Disallow: /
User-agent: Meta-ExternalAgent
Disallow: /
User-agent: Amazonbot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /
# AI search crawlers
User-agent: OAI-SearchBot
Disallow: /
User-agent: Claude-SearchBot
Disallow: /
User-agent: PerplexityBot
Disallow: /
User-agent: Meta-WebIndexer
Disallow: /
User-agent: Amzn-SearchBot
Disallow: /
User-agent: MistralAI-Index
Disallow: /
User-agent: DuckAssistBot
Disallow: /
# User-triggered AI fetchers
User-agent: ChatGPT-User
Disallow: /
User-agent: Claude-User
Disallow: /
User-agent: Perplexity-User
Disallow: /
User-agent: Meta-ExternalFetcher
Disallow: /
User-agent: Amzn-User
Disallow: /
User-agent: MistralAI-User
Disallow: /This costs the most. It asks every AI search crawler named here to leave the store out of its search, and every fetcher to leave its pages unread when a shopper asks about a product. It does not touch Googlebot or bingbot, which Google's AI Overviews and AI Mode and Bing's Copilot answers draw on. Several fetchers say robots.txt may not apply to them, and Google's user-triggered fetchers have no token to block.
The snippet names every crawler listed above as a training crawler, training control, AI search crawler or user-triggered fetcher, and no other: not the search engines, and not Google-CloudVertexBot, which Google says crawls at a site owner's request to build Vertex AI Agents and has no effect on Google Search. Cloudflare's managed robots.txt, which Cloudflare says tells known AI crawlers to stay away, also disallows Bytespider and CCBot. Common Crawl says CCBot checks robots.txt first and fetches a page only if crawling it is allowed, and its own instruction for keeping CCBot out is User-agent: CCBot followed by Disallow: /.
Do not block AI crawlers with User-agent: * and Disallow: /. That keeps out every crawler that honours robots.txt and follows your * group: Googlebot, bingbot and Storebot-Google, which feeds Google Shopping, among them, wherever no other group names them, and Applebot too unless your file names it or Googlebot. So it shuts search engines out as well.
Where to make the change on Shopify, WooCommerce and BigCommerce
Shopify
Shopify writes a default robots.txt for every store, at the root of its primary domain, and changes go through a theme template, robots.txt.liquid, which replaces the generated file. First open your store's /robots.txt and save a copy, as Shopify recommends, to compare with later. Then, from your Shopify admin, go to Online Store; for your published theme, open the actions menu and choose Edit code; click Add a new template, select robots, and click Create template. Shopify's help then has you make your changes to the default template and save them, and recommends keeping its Liquid rather than replacing it with plain text, so that Shopify can keep the file up to date; Shopify's developer page says a rule that is not part of a default group goes outside the Liquid that prints the default rules. Add the groups for your choice on new lines above that Liquid, at the very top of the template, with a blank line between your last line and the Liquid, and save. Placed first, nothing the Liquid prints can run into your first group (see the paste rule above). A template that is empty prints none of Shopify's rules; if yours is, first copy in the example on Shopify's developer page for the robots.txt.liquid template, which prints the default rule groups.
Then reopen /robots.txt to confirm the change, and compare it with your copy. Your groups should come first, as you pasted them, with User-agent: * starting its own line after your last Disallow: /. Shopify notes that a template printing its default rule groups can carry rules its current default file no longer has, such as Disallow: /search and Disallow: /policies/, and the second, where it sits in your * group, keeps every crawler that follows that group, Googlebot included, out of your store's /policies/ pages. If any Disallow: line containing /policies/ or /search appears where your copy had none, take out each one with the Liquid on Shopify's developer page under "Remove a default rule from an existing group": its example skips one exact rule, /policies/, so add a condition for each line you find. Or delete the template to go back to Shopify's default file, which takes your groups out with it. Either way, reopen /robots.txt afterwards and compare it with your copy again to confirm the change.
Shopify's help is direct about the risks. The template is an unsupported customization that its support team cannot help with, and incorrect use can lose a store all of its traffic. It asks you to remove older workarounds, such as a third-party service like Cloudflare, before you edit. Changes are instant, though crawlers do not always react at once. The template lives in the theme, and Shopify notes that uploading a theme from the Themes section of the admin does not import robots.txt.liquid, so check the file again after you change themes.
Shopify says two more things about AI crawlers. Product data for agentic storefronts you have turned on, such as ChatGPT or Microsoft Copilot, reaches them through Shopify Catalog, independently of robots.txt, so a block in robots.txt or at the network affects only what crawlers read on the open web. And bot management at the network layer is handled for stores on Shopify with no action needed; Shopify does not recommend a proxy in front of a Shopify store.
WooCommerce and WordPress
WordPress generates robots.txt in code: one * group that keeps crawlers out of the admin area, to which WooCommerce adds rules for its add-to-cart links and its log and upload folders. Changes go through WordPress's robots_txt filter, so they are made in code or with a plugin that edits robots.txt. Add your groups after the * group, never inside it: a User-agent: line that follows a group's rules ends that group, so any * rules below your lines would belong to your new group instead.
A site hosted on WordPress.com has a setting for AI training and third-party use: Settings, then Reading, then the Prevent third-party sharing box under Public. WordPress.com says that ticking it keeps the site's public content out of its network of third-party content and research partners, and adds known AI bots to the disallow list in the site's robots.txt, though it is up to AI platforms to honour that. It does not say which bots it adds, so the setting may keep AI search crawlers and fetchers out as well as training crawlers; open the live file after you change the setting to see which bots it names.
BigCommerce
Go to Settings, then Website, and scroll to the Search Engine Robots section; your user account needs the Manage Settings permission, and with Multi-Storefront each storefront has its own file. Paste your groups into the box after whatever groups it already holds, never inside one, and save. BigCommerce combines what you save with its system defaults, merging the rules into a single entry when the same crawler appears in both.
It also applies rules of its own to a list of known high-volume and AI crawlers: a 10-second crawl delay, and blocks on sensitive pages such as cart, checkout, account, admin and filter pages. You cannot see or edit these in settings; they appear only in your live robots.txt, as a group that names each crawler on the list, and a crawler named there follows that group, not your * group. In the file BigCommerce serves for its own demo store, that group names many of the crawlers in this guide, among them GPTBot, ClaudeBot, OAI-SearchBot, PerplexityBot and ChatGPT-User, and Applebot as well. BigCommerce says no entries are needed to allow AI bots.
BigCommerce's help says that when the same crawler appears in your rules and its defaults, the rules are merged into a single entry. It does not say what that does for a crawler on its shared list, where a Disallow: / merged into the shared group would keep out every crawler the group names: Applebot, the AI search crawlers OAI-SearchBot, PerplexityBot and Meta-WebIndexer, and facebookexternalhit, which Meta says crawls pages shared on Facebook, Instagram or Messenger, among them. So after you save, open your live robots.txt and check that each Disallow: / sits in a group that names only crawlers you chose to keep out. If one does not, take the groups you added back out, save, and check the file again. Where BigCommerce merges your block into its shared group, the block applies to every crawler that group names, so robots.txt cannot keep that one crawler out on its own; BigCommerce's help points stores that need more robust controls than robots.txt to a service such as Cloudflare.
What robots.txt does not do
- It asks; it does not enforce. RFC 9309 says its rules are not a form of access authorization, Shopify calls them directional and advisory, and Cloudflare says compliance is voluntary. Listing a path in robots.txt also makes it public.
- Fetchers acting for a user may not apply it. OpenAI says robots.txt rules may not apply to ChatGPT-User's visits, because a user initiated them; Perplexity says Perplexity-User generally ignores them; Google says its user-triggered fetchers generally ignore them; Meta says Meta-ExternalFetcher may bypass them; and Amazon says Amzn-User may not follow all of them. Anthropic's statement that its bots honour robots.txt makes no exception for Claude-User.
- It takes time. Crawlers keep a copy of robots.txt: OpenAI gives about 24 hours for its search results to adjust, Amazon about 24 hours, Perplexity and Meta up to 24 hours, and DuckDuckGo 72 hours. Amazon also says it may use a cached copy from the last 30 days.
- It covers one host. Each host serves its own
/robots.txt, sowww.example.comandshop.example.comare separate files, and Anthropic asks site owners to add its block for every subdomain they want to opt out. - It controls crawling, not indexing. Google notes that blocking Googlebot from a page does not stop the page's address appearing in search results, and Bing's guidelines say robots.txt controls crawl access, not indexing.
Crawl-delayis not part of the standard. Anthropic and Common Crawl say they honour it; Google, Apple and Amazon say they do not.- A name in your logs can be borrowed. Anyone can send a crawler's user-agent string, and Common Crawl warns that some crawlers falsely identify themselves as its CCBot. Most of the operators here publish lists of the IP addresses their crawlers use, so a visit can be checked against them.
A CDN or firewall can decide first
robots.txt is not the only gate. A CDN or bot-protection service in front of a store can turn crawlers away before they read a page, and Google's best practices for its AI features include making sure crawling is allowed "in robots.txt, and by any CDN or hosting infrastructure".
Cloudflare is a common example. In July 2025 it said every new domain would be asked at sign-up whether to allow AI crawlers, and its AI bot policies (Security Settings, then Configure AI bot policies) can block AI crawlers at the network, whatever robots.txt says. Its documentation, last updated July 1, 2026, said that on September 15, 2026 it would set new defaults for new domains: crawlers it classes as Training, and agents acting for a person in real time such as chat fetch bots, blocked on pages that display ads, with search crawlers still allowed; and crawlers used for both search and training blocked by every setting that blocks AI training. Its managed robots.txt setting acts on the file instead: it puts Cloudflare's own disallow rules for known AI crawlers at the top of your robots.txt. Cloudflare's example of that block carries a Content-signal line, which Cloudflare defines and RFC 9309 does not, and disallows Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and Meta-ExternalAgent. If your domain is on Cloudflare, check both settings.
BigCommerce's help points stores that need stronger controls than robots.txt to a service such as Cloudflare. Shopify handles bot management itself, as above.
Controls that live on the page
Some uses are set page by page in the page's own HTML, with robots meta tags (and, for Google and Bing, the data-nosnippet attribute), rather than in robots.txt. A tag written as <meta name="robots"> speaks to every crawler that honours it, so one meant for one company's AI answers reaches the others too. Apple, Bing and Google each document a form addressed to their own crawler alone: <meta name="applebot">, <meta name="bingbot"> and <meta name="googlebot">.
- Google:
nosnippet,data-nosnippet,max-snippetandnoindexlimit what Search shows from a page, AI features included. They apply to all of Google Search, not only its AI features: Google saysnoindexkeeps the page out of its search results, andnosnippetremoves the page's text snippet from every form of search result, AI Overviews and AI Mode among them, and keeps its content from being used as a direct input for AI Overviews and AI Mode. - Bing:
noarchivekeeps a page's content out of Copilot answers and grounding results,nocachelimits Copilot to the page's address, title and snippet, and Bing saysnosnippetanddata-nosnippetstop it showing captions and may limit the quality of Copilot's citations. Bing said in September 2023 that it would not use content labeled NOARCHIVE to train Microsoft's generative AI foundation models, that content with either tag still appears in its search results, and that content with both is treated as NOCACHE. - Amazon: a page's
noarchiverobots meta tag means "do not use the page for model training". Amazon's page documents robots meta tags, not a form for its own crawlers, so the documented tag is<meta name="robots" content="noarchive">, which Bing reads too: it also keeps the page out of Copilot answers. - Apple: the
nosnippetmeta tag keeps content out of the broad world knowledge answers Apple's AI models give. Apple says it also stops Applebot generating a description or web answer for the page, so Apple's suggestions to visit the page show only its title.
Read your live robots.txt
- Open
/robots.txton your store's domain in a browser, for every host you sell from. That is the file crawlers read, with your platform's rules and any CDN's additions in it. - Find the
User-agent: *group, then any group that names a token from this guide. A crawler obeys the groups that name it, and only those; every other crawler obeys*, apart from the Applebot and Amzn-SearchBot fallbacks above and Google's AdsBot, which Google says ignores the*group with the ad publisher's permission. - Check that each group meant as a block says
Disallow: /and names only crawlers you chose to keep out, and that a group naming a crawler you want in still keeps it out of carts and checkouts. Check too that the group just above each block has anAlloworDisallowrule of its own, or that the block opens the file, and that no other group that names a crawler you are keeping out has anAllowline. - If the domain sits behind Cloudflare or another CDN, check its bot settings too: robots.txt cannot show what they block.
A Findwise scan fetches your robots.txt and works out which groups Googlebot and bingbot would obey, and the AI search crawlers OAI-SearchBot, Claude-SearchBot and PerplexityBot, and the user-triggered fetchers ChatGPT-User, Claude-User, Perplexity-User and Meta-ExternalFetcher. It reads one case differently from the standard: RFC 9309 and Google's parser join a group that has a Crawl-delay: line but no Allow or Disallow rule to the group below it, and the scan reads it as a group of its own. Its crawlability score counts whether Googlebot, those seven and the scan's own crawler may read your product and collection pages, and your policy, help and guide pages. A group naming any other search crawler the scan knows, Applebot and DuckDuckBot among them, counts the way those seven do. Training crawlers are not counted, so a block on training alone does not lower it.
The scan reports a file whose groups name at least one crawler of two or more of the three kinds and keep it out of every page, counting those other search crawlers as AI search: one training crawler and one fetcher are enough. A crawler kept out only by your * group is not reported this way, though if the score counts it, the score still falls. It goes by the crawlers it knows by name, and it does not yet know Meta-WebIndexer, Amzn-SearchBot, MistralAI-Index, DuckAssistBot, MistralAI-Training, Amzn-User or MistralAI-User. It also reports Googlebot kept out of product or collection pages, and a robots.txt that does not answer, is refused, fails with a server error or does not parse. Each finding cites what it read: the groups in your file, or the answer the request for it got. The scan's own crawler is FindwiseBot, and section 3 of Findwise's crawl policy says it follows "the rules for FindwiseBot if there are any, otherwise the rules for *" (how it identifies itself, and how to refuse it). Findwise estimates from public pages; it does not test AI assistants and cannot say what any assistant shows for your store.
Check your own store
A Findwise scan reads a sample of your store's public pages and reports where the product facts, structure and policies are thin for AI-assisted shopping, each finding with the page it was seen on.
Sources
- IETF RFC 9309: Robots Exclusion Protocol
- Google for Developers: How Google interprets the robots.txt specification
- OpenAI: Overview of OpenAI crawlers
- Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?
- Google for Developers: Google's common crawlers
- Google for Developers: Google's user-triggered fetchers
- Google for Developers: Google crawler (user agent) overview
- Google Search Central: Googlebot
- Google Search Central: AI features and your website
- Google Search Central: Robots meta tag, data-nosnippet and X-Robots-Tag specifications
- Perplexity: Perplexity crawlers
- Apple Support: About Applebot
- Bing Webmaster Tools: Which crawlers does Bing use?
- Bing Webmaster Tools: Bing Webmaster Guidelines
- Bing Webmaster Blog: Announcing new options for webmasters to control usage of their content in Bing Chat (September 2023)
- Meta for Developers: Meta web crawlers
- Amazon: Amazonbot
- Mistral AI docs: Mistral's crawlers and robots.txt
- DuckDuckGo Help: DuckAssistBot
- Common Crawl: CCBot
- Common Crawl: FAQ
- Shopify Help Center: Editing robots.txt.liquid
- Shopify developer docs: Customize robots.txt
- Shopify developer docs: robots.txt.liquid
- WordPress developer reference: do_robots()
- WordPress developer reference: the robots_txt filter
- WooCommerce source code: class-woocommerce.php (its robots.txt additions)
- WordPress.com Support: Make your website public (Prevent third-party sharing)
- BigCommerce Help Center: Understanding the robots.txt file
- BigCommerce: the robots.txt of its Cornerstone theme demo store, as served on October 7, 2026
- Cloudflare docs: Managed robots.txt
- Cloudflare docs: Block AI bots
- Cloudflare press release: Cloudflare just changed how AI crawlers scrape the Internet at large (July 1, 2025)
