Free tool

Free AI crawler checker for law firm websites

Enter your domain below. The check fetches your robots.txt, reads it the way a crawler does, and tells you which AI bots your site turns away. It shows the raw file too, so you can read the lines yourself rather than trust our reading of them. Thirty seconds, no form. This is the automated version of cause 4 in our longer diagnostic, why your law firm isn't showing up in ChatGPT.

We fetch /robots.txt from the domain you enter and evaluate it against the homepage path /, the way RFC 9309 tells a crawler to. Nothing is stored, and no email is asked for.

Enter a domain to see which of the ten AI crawlers below your robots.txt turns away. A competitor's domain works too — the file is public on every website.

The AI crawler registry

Most advice on this subject goes wrong within months, because the bot list changes and nobody updates the post. This table is dated and links to each operator's own documentation, so you can check any row yourself. It is as current as the date on it — the only freshness claim we make. Verified on 2026-08-17.

User agentOperatorWhat it doesIf you block it
GPTBotOpenAITraining crawl for foundation models docsContent excluded from training. Does not remove you from ChatGPT's search answers.
OAI-SearchBotOpenAISurfaces sites in ChatGPT's search features docsOpenAI states opted-out sites "will not be shown in ChatGPT search answers, though can still appear as navigational links," and notes ~24 hours for a robots.txt change to take effect.
ChatGPT-UserOpenAIFetches a page when a user asks about it docsOpenAI states robots.txt rules "may not apply" — the fetch is user-initiated, not automatic crawling.
PerplexityBotPerplexityIndexes pages to surface and link in results; not used for training docsYou won't be surfaced in Perplexity results.
Perplexity-UserPerplexityVisits pages in response to a user's question docsPerplexity states it "generally ignores robots.txt rules" because a user initiated the request.
GooglebotGoogleGoogle's Search crawler; AI Overviews and AI Mode are part of Search docsYou leave Google Search entirely. Almost never the right move.
Google-ExtendedGoogleWhether crawled content trains Gemini and grounds Gemini Apps and Vertex AI docsGoogle states it "does not impact a site's inclusion in Google Search" and is not a ranking signal — so this does not remove you from AI Overviews.
ClaudeBotAnthropicTraining crawl docsFuture content excluded from training datasets.
Claude-SearchBotAnthropicIndexes content for Claude's search results docsVisibility in Claude's search results may drop.
Claude-UserAnthropicVisits pages when a Claude user asks a question docsClaude won't fetch the page mid-conversation. Anthropic states its bots — this one included — honor robots.txt.

Two rows are the reason this page exists. Blocking GPTBot does not hide you from ChatGPT's search answers, and blocking Google-Extended does not hide you from AI Overviews. Both are training-and-grounding controls; the bots that decide whether you can be quoted in a live answer are different bots. A firm that “blocked ChatGPT” in 2023 by disallowing GPTBot can still be fully visible in ChatGPT's search answers today.

What this check cannot see

robots.txt is one of at least four places your site can refuse a crawler. The check reads the first one.

  1. Your CDN or firewall. Bot-management rules block by user agent or IP before the request reaches your server, and robots.txt says nothing about them. It is also the block least likely to be visible to you: on 1 July 2025 Cloudflare began asking every newly signed-up domain whether to allow AI crawlers (press release). If your site was migrated or rebuilt after that date, someone answered that question on your behalf — and the answer lives in a dashboard, not in robots.txt.
  2. Server-side rules. A security plugin, a WAF ruleset, or a hand-written nginx rule can return a 403 to a named user agent while robots.txt reads as wide open.
  3. Rate limiting. A crawler throttled to a trickle looks allowed and behaves blocked.
  4. Snippet-level controls. Google points publishers to nosnippet, data-nosnippet, max-snippet and noindex to limit what Search shows, and treats AI Overviews and AI Mode as part of Search (AI features and your website). A data-nosnippet wrapper added to your bio pages years ago is invisible in robots.txt.

Two more limits are worth stating plainly. The check evaluates your file against the homepage path /, so a rule that blocks only a sub-path — a client portal, say — correctly reports as no block on the site as a whole. And robots.txt is a request, not a lock: the standard that defines it says its rules are not a form of access authorization. Compliance is voluntary — two agents in the registry above say, in their operators' own words, that they may not follow it, and the tool labels those rows rather than pretending a block there is binding.

For the longer version — why a block that looked harmless in 2023 may still be costing you answers today, and what each crawler actually controls — read the guide to AI crawler access for law firm sites.

What a clean result does not mean

A clean result means your robots.txt is not turning crawlers away. It does not mean an assistant is reading you, and it does not mean an assistant recommends you. Crawler access is a precondition, not a cause.

The measurement on the other side of that line is harder and less certain, and we would rather say so than sell you a straight line. AI answers are probabilistic: the same question asked twice does not reliably return the same firms. One check is a snapshot of one moment. That is why we report recommendation data as a trend across many questions, many runs and several assistants rather than as a score, and why our methodology is published rather than proprietary.

See whether the assistants that can reach you actually recommend you — free

A crawler that can reach you is not the same as an assistant that recommends you. The free Briefly check asks AI assistants the hiring questions real clients ask in your practice area and market, then shows whether your firm is recommended, which firms are named when it isn't, and which sources shaped the answer.

Check your firm free →

No signup. No credit card. No sales call — you see the actual answers.

John Rice builds and operates the scan engine behind Briefly, which runs client-style lawyer-hiring questions across ChatGPT, Gemini, Perplexity, Google AI Overviews, and Google AI Mode on a recurring schedule and stores every answer it collects. He is not a lawyer; he measures what AI assistants say, with receipts. The crawler behaviours described above come from each operator's own published documentation, linked in every row and re-verified on the date shown. Methodology · About John

Briefly measures what AI assistants say. It does not rank, rate, or endorse attorneys, and nothing here is legal advice.