Blog

Is your law firm's website blocking AI assistants?

John Rice5 min read

Run the free AI crawler check on your domain. It fetches your robots.txt, reads it the way a crawler does, and gives one verdict per AI crawler — allowed, blocked, or unknown — with the line that decided it, quoted and numbered. It shows the raw file too. A second or two, no form. This is the automated version of cause 4 in why your law firm isn't showing up in ChatGPT.

How does a law firm find out whether its website is blocking AI assistants? Open yourfirm.com/robots.txt and look for a Disallow: / line under any AI crawler's name — GPTBot, OAI-SearchBot, PerplexityBot, Google-Extended and ClaudeBot are the common ones.

Two things about the verdicts. They are worked out for your homepage, so a block scoped to one path — a client portal, say — reads as allowed. And if your server times out, errors, or refuses us the file, every verdict comes back unknown rather than allowed.

The AI crawler registry

The bot list changes, so undated advice goes stale fast. Every row links to the operator's own documentation. Verified on 2026-08-17.

User agentOperatorWhat it doesIf you block it
GPTBotOpenAITraining crawl (docs)Content kept out of training. Not out of ChatGPT's search answers.
OAI-SearchBotOpenAISurfaces sites in ChatGPT's search (docs)OpenAI states opted-out sites "will not be shown in ChatGPT search answers, though can still appear as navigational links."
ChatGPT-UserOpenAIFetches a page a user asks about (docs)OpenAI states robots.txt rules "may not apply" — the fetch is user-initiated.
PerplexityBotPerplexityIndexes pages to link in results (docs)You won't be surfaced in Perplexity results.
Perplexity-UserPerplexityVisits a page a user asks about (docs)Perplexity states it "generally ignores robots.txt rules" for user-initiated requests.
GooglebotGoogleSearch crawler; AI Overviews and AI Mode are part of Search (docs)You leave Google Search entirely. Almost never the right move.
Google-ExtendedGoogleWhether your content trains Gemini and grounds Gemini Apps (docs)Google states it "does not impact a site's inclusion in Google Search" — so it is not a way out of AI Overviews.
ClaudeBotAnthropicTraining crawl (docs)Future content excluded from training datasets.
Claude-SearchBotAnthropicIndexes content for Claude's search results (docs)Visibility in Claude's search results may drop.
Claude-UserAnthropicVisits a page a Claude user asks about (docs)Claude won't fetch it mid-chat. Anthropic states its bots honor robots.txt.

Two rows are the reason this page exists. Blocking GPTBot does not hide you from ChatGPT's search answers, and blocking Google-Extended does not hide you from AI Overviews. Both are training controls; the bots that decide whether you can be quoted in a live answer are different bots. Our own reading of those docs — Briefly's, not a statement from OpenAI — is that a firm which "blocked ChatGPT" in 2023 by disallowing GPTBot may well still be visible in ChatGPT's search answers today. Check your file for both names.

What this check cannot see

robots.txt is one of several places your site can refuse a crawler, and the only one this check reads.

  1. Your CDN or firewall. Bot rules block by user agent or IP before the request reaches your server. Since 1 July 2025, Cloudflare has asked every newly signed-up domain whether to allow AI crawlers (press release) — an answer that lives in a dashboard, not in robots.txt.
  2. Server-side rules. A security plugin or an nginx rule can return a 403 to one named user agent while robots.txt reads as wide open.
  3. Rate limiting. A crawler throttled to a trickle looks allowed and behaves blocked.
  4. Snippet-level controls. Google points publishers to nosnippet, data-nosnippet, max-snippet and noindex to limit what Search shows, and treats AI Overviews and AI Mode as part of Search (docs). As of 17 August 2026, that documentation describes no control that keeps ordinary search snippets while removing you from AI Overviews.

And robots.txt is a request, not a lock. The standard says its rules "are not a form of access authorization" (RFC 9309). Two agents above say, in their operators' own words, that they may not follow it.

What to do with a block you find

Say the check reports Disallow: / under GPTBot and PerplexityBot for Bellweather & Sunde Family Law (a fictional firm). Three steps:

  1. Find out who added it and when. Usually a web vendor, an SEO plugin default, or a 2023-era decision to stay out of AI training. Reasonable then; worth re-taking now.
  2. Decide per bot, not in bulk. A firm can decline the training crawlers and still allow the search bots.
  3. Scope any block narrowly. Block the client portal path, not the whole site.

Then ask your web vendor the question robots.txt cannot answer: "are we actually serving pages to OAI-SearchBot and PerplexityBot, or is the CDN dropping them?" Server logs settle it.

FAQ

Should my law firm block GPTBot? Blocking it keeps your content out of OpenAI's model training. It does not remove you from ChatGPT's search answers — OAI-SearchBot governs that. If the goal is to be recommended, allow the search bots.

Will allowing AI crawlers hurt my Google rankings? No documented penalty either way. Google states that Google-Extended "does not impact a site's inclusion in Google Search" and is not a ranking signal (docs).

The check says I'm not blocking anything. Why am I still not mentioned? Crawler access is a precondition, not a cause. The usual reasons are thin directory profiles, little third-party coverage, and firm details that differ from listing to listing — the six-cause diagnostic covers each one.

How current is this list? As current as the date printed on it, which is the only freshness claim we make. John Rice re-reads each operator's documentation and moves that date when he does. Every row links to its source, so you can check the one that matters to you.

John Rice builds and operates the scan engine behind Briefly, which runs client-style lawyer-hiring questions across ChatGPT, Gemini, Perplexity and Google's AI surfaces on a recurring schedule. He is not a lawyer; he measures what AI assistants say, with receipts. Methodology · About John

Briefly measures what AI assistants say. It does not rank, rate, or endorse attorneys, and nothing here is legal advice.

Sources