FindwiseBot and methodology
FindwiseBot is the crawler behind the Findwise store scan. It reads a small sample of a store's public pages so the scan can report on what AI shopping systems would find there. The short answers are here; the full policy follows.
The exact user agent
FindwiseBot/0.1 (+https://getfindwise.com/bot)How to stop it
Add this to your robots.txt. FindwiseBot answers to the FindwiseBot group, and applies the * group when there is no group of its own.
User-agent: FindwiseBot
Disallow: /It stops on the next scan, and blocking needs no message to us. A blocked scan still runs — it reports on what it could reach and says on the report that coverage was limited. There is no mode that crawls around a robots rule.
The limits it keeps
- What it fetches
- Public store pages only, and a few pages of the store's own help centre when the store's pages link it (see Scope). No logins, no forms, no checkout. On stores built on a common platform it may also request that platform's standard policy addresses (for example /policies/refund-policy on Shopify), robots.txt permitting.
- Signed identity
- When configured, Ed25519 HTTP message signatures identify FindwiseBot as https://www.getfindwise.com. Public keys are at /.well-known/http-message-signatures-directory on that hostname. Without a key, requests remain unsigned and no directory is published. Signing does not mean verification or guarantee access; robots and crawl limits still apply.
- Pages per scan
- At most 30.
- Request rate
- One request per second per host, two at a time at most.
- Page size
- 1 MB per page. Anything larger is truncated.
- Time budget
- 90 seconds for a whole scan, 10 seconds per page.
- Crawl-delay
- Honoured, up to 10 seconds. A longer delay is treated as 10 seconds.
- Rate limits
- A 429 Too Many Requests is honoured. If a store turns away a free scan, the one we run when someone enters the store on the Findwise site, with nothing but 429s, so no page is read, the scan stops and may be tried once more later, as a new scan under these same limits: after the wait the store's Retry-After asks for, but no sooner than 5 minutes, or after 10 minutes when it names none. A store that asks for more than 30 minutes is not asked again, a try that cannot start within 15 minutes of its time is not made, it never runs beside another scan of the same store, and there is never a second one.
- Redirects
- At most 5 hops. The address you submit may be followed once to another public https site. A further hop is not followed.
- Scope
- One site at a time. www and the bare domain count as the same site. Other subdomains are not followed, with two exceptions. First, when the store's own pages link its FAQ or help centre on the store's own domain or another address under it (for example help.brand.com for brand.com, or brand.com for shop.brand.com), the scan may read up to 2 pages there (3 in a deep scan), counted inside the page limit, after reading that address's robots.txt. If that robots.txt cannot be read, nothing there is read. A help-desk provider's own domain is never read. Second, the address you submit: if its first response redirects once to another public https address, including a subdomain of itself or a different site, the scan follows that one hop, reads that host's robots.txt first, and the report names both addresses. A second hop is not followed.
Findwise methodology
How does Findwise score a store?
Versioned automated checks use evidence from a sample of public pages to produce the scores, grade, priorities and findings. Each finding points to the evidence behind it. AI wording is optional and does not change the scores or decide which gaps matter most.
What does Estimated coverage mean?
Estimated coverage compares the captured store evidence with a versioned set of buyer-question patterns. It describes how much supporting information the scan found. Findwise does not test AI assistants directly and does not predict traffic or sales lift.
Where does the example report come from?
The example uses a sample store built from test pages. It is a stored scan produced by the same scanner and read through the same report service as a visitor's scan, not a live customer scan or a report written for display.
Does a scan cover every page?
A scan reads a bounded sample of publicly reachable pages and honours robots.txt. Missing or blocked pages limit what the report can conclude. The crawler limits and full policy on this page explain that boundary.
Findwise Bot And Crawl Policy
Last updated October 3, 2026 · Version bot-2026-10-03-v0.7
This policy explains how FindwiseBot is designed to crawl public ecommerce websites for Findwise scans and reports.
Contact: legal@getfindwise.com
1. What FindwiseBot Does
FindwiseBot fetches public ecommerce pages to help Findwise evaluate whether AI shopping systems can find, understand, compare, and cite a store's product information.
A FindwiseBot visit does not mean the website owner requested, endorsed, approved, purchased, or is affiliated with Findwise.
FindwiseBot may analyze:
- Home pages.
- Product detail pages.
- Collection and category pages.
- FAQ, support, help, shipping, returns, warranty, and policy pages, including a few pages of the store's own help centre when the store's pages link it (see section 4).
- Review, comparison, guide, alternative, and buyer-intent pages.
- Public robots.txt and sitemap files.
- Public metadata, schema, headings, internal links, and visible page text.
FindwiseBot does not intentionally crawl authenticated pages, carts, checkout pages, account pages, admin pages, private APIs, destructive action URLs, or non-public content. It refuses these by path — cart, checkout, account, login, admin, wp-admin, api, graphql and similar — and by query parameter, so a link that performs an action on request rather than showing a page (?add-to-cart=, ?remove_item=, ?logout=) is never followed. If a public page contains information that should not have been public, contact legal@getfindwise.com and secure the page at the source.
2. User Agent
Production crawler user agent:
FindwiseBot/0.1 (+https://getfindwise.com/bot)The version number may change.
When configured for its final production identity, FindwiseBot also identifies requests with Ed25519 HTTP message signatures: Signature-Input, Signature and Signature-Agent. The identity is https://www.getfindwise.com, with its public keys at https://www.getfindwise.com/.well-known/http-message-signatures-directory. The directory contains public keys only and signs its response to bind the keys to that hostname. Without a signing key, requests remain unsigned and the directory is not published. Signing alone grants no special access and does not guarantee a store will admit a request. Registration of this final identity follows production domain cutover; no temporary staging identity is registered. Signing changes none of the robots, rate, page or access limits below.
3. Robots.txt
FindwiseBot is designed to respect robots.txt in production.
FindwiseBot will first look for rules that specifically apply to FindwiseBot. If none exist, it will follow rules for *.
Before it requests or stores an address, FindwiseBot removes query parameters whose names mark credentials (token, key, session, signature and similar), and it checks robots.txt against the address it will actually request.
Example:
User-agent: FindwiseBot
Disallow: /Findwise may also honor reasonable requests sent to legal@getfindwise.com. Findwise may verify domain ownership or authority before applying a manual block, correction, or removal.
If robots.txt returns a server error, times out, or cannot be read, FindwiseBot continues at its normal polite rate with no rules applied, and records the failure on the report as a limitation that lowers the scan's confidence. It does not treat an unreadable file as permission to crawl harder, and it does not silently drop the scan.
The store's own help centre on another address (section 4) is treated more strictly. FindwiseBot reads that address's robots.txt before requesting any page there, and follows its FindwiseBot rules, else its * rules. If that address's robots.txt returns a server error, times out, redirects to another host, or cannot be read, FindwiseBot requests nothing else there, and the report says those pages were not read. A missing file (404 or 410) means no rules there.
Robots.txt is not a security boundary. Website owners should not place private or sensitive information on public URLs, even if those URLs are disallowed in robots.txt.
4. Crawl Limits
FindwiseBot is designed to crawl politely:
- Same-origin pages only by default. Other subdomains are not followed, with two exceptions.
- The store's own FAQ or help centre. When the store's own pages link its FAQ or help centre on the store's own domain or another address under it (for example
help.brand.comforbrand.com, orbrand.comforshop.brand.com), the scan may read up to 2 pages there (3 in a deep scan), counted inside the page limit, after reading that address's robots.txt (section 3). One such address per scan, reached only through a link whose text or address names its FAQ or help centre, or through two of its shipping, returns or warranty links. Only that address: a redirect from it to any other host is not followed, and no sitemap there is requested. A help-desk provider's own domain (for examplebrand.gorgias.help) is never read, and neither is an address whose name says it is a shop, a checkout, an account area, a returns portal, a blog or a community. - The address submitted for the scan, and only on its first response. If that address redirects once to another public https address — a subdomain of itself (for example
brand.comtoshop.brand.com) or a different site (for exampleoldname.comtonewname.com) — the scan follows that one hop, reads the destination's robots.txt before requesting anything, and the report names both addresses and is about the destination. A second hop to a further site is not followed. An http destination is not followed. The address that was submitted stays the one the scan is filed under. - Public HTML pages only.
- On stores built on a common platform it may also request that platform's standard policy addresses (for example
/policies/refund-policyon Shopify), robots.txt permitting. - Static assets ignored.
- Page count limited by scan mode.
- Request rate and concurrency limited per host.
- Page-size and timeout caps.
- Crawl-delay honored, up to a cap of 10 seconds. A longer delay is treated as 10 seconds.
- Fetch failures recorded as facts, not bypassed.
- One later try after a rate limit on a free scan. A free scan is the one we run when someone enters a store on the Findwise site. If a store turns a free scan away only with
429 Too Many Requests, so that no page is read, the scan stops, and FindwiseBot may try that scan once more, as a new scan under every limit on this page: after the wait the store'sRetry-Afterheader asks for, but no sooner than 5 minutes, or after 10 minutes when it names no wait. A store that asks for longer than 30 minutes is not asked again, a later try that cannot start within 15 minutes of its time is not made, and it never runs while another scan of the same store is running. There is never a second one. - Further attempts for a paid report. When the scan for a full report or a rescan someone has bought cannot finish, FindwiseBot may try that scan again, up to 3 attempts in all: the second about 10 minutes after the first starts, and the third about 30 minutes after the second. No attempt starts later than 2 hours after the payment, except that a rescan Findwise itself held back, for example while scanning was paused, starts once it can. FindwiseBot does not try again when robots.txt blocks the scan, when the site is not a store, or when the store's bot protection turns it away. Each attempt is a new scan under every limit on this page, and it never runs while another scan of the same store is running.
Initial production defaults, subject to lower limits or refusal where operationally appropriate:
- Quick Scan: up to 8 HTML pages, at most 2 of them on the store's own help centre.
- Deep Preview / paid scan: up to 30 HTML pages, at most 3 of them on the store's own help centre.
- Maximum page size: 1 MB of HTML per page. A larger page is truncated, and the report says so.
- Per-page timeout: 10 seconds.
- Host request rate: approximately 1 request per second, on the store and on its help centre alike.
- Concurrency: no more than 2 requests per host.
These limits may be adjusted to protect site owners, reduce abuse, improve reliability, or support paid scans. Findwise should not publish more specific limits than the production crawler can enforce and log.
If the implementation changes, this page should be updated before the change is used in production. Site owners should rely on the current published policy and their own server logs rather than assuming any single scan reflects all future behavior.
Findwise may stop, slow, deny, or block a scan if we detect unusual load, abuse risk, legal risk, security risk, inaccurate configuration, domain-owner objection, or other operational concern.
5. What Findwise Does Not Do
FindwiseBot does not:
- Bypass login, paywalls, robots.txt, CAPTCHAs, or technical access controls.
- Submit forms, add items to cart, initiate checkout, or perform destructive actions.
- Crawl private customer accounts or admin systems.
- Intentionally collect payment card numbers, passwords, API keys, protected health information, government identifiers, or other sensitive personal information.
- Use public website scans as a people-search, credit, employment, insurance, housing, or regulated decisioning service.
- Sell personal information collected through scans.
- Claim that its scan is complete, exhaustive, or equivalent to how any particular AI provider, search engine, or shopping agent will read the site.
- Guarantee that crawled content is accurate, lawful, endorsed, current, complete, or suitable for any customer's purpose.
6. Raw Crawl Data And Reports
Findwise may store raw public HTML, extracted evidence, scan metadata, findings, and reports to provide the Services, debug scan quality, support customers, and improve scoring.
Raw crawl data is not exposed in public free reports. Paid reports may include selected evidence snippets, page references, and recommendations. Findwise may remove, redact, or withhold evidence where privacy, security, legal, or third-party concerns apply.
Free preview report links should be unguessable and should not be indexed by search engines unless Findwise intentionally publishes the report as an example, benchmark, or marketing page after review.
7. How To Block Or Contact Findwise
To block FindwiseBot with robots.txt:
User-agent: FindwiseBot
Disallow: /For questions, removal requests, crawl issues, or abuse reports, contact legal@getfindwise.com and include:
- Your domain.
- The issue or request.
- Relevant log lines if available.
- A contact email for follow-up.
Findwise may verify requests before making changes.
