Blog

Does `max-snippet` stop AI Overviews?

John Rice24 min read
<!-- PUBLISHER NOTES. Three of these are corrections owed to pages outside this file. Do not edit those files from this draft; another session owns them. 1. CORRECTION REQUIRED — /blog/ai-crawler-check, section "What this check cannot see", item 4, currently reads: "As of 17 August 2026, that documentation describes no control that keeps ordinary search snippets while removing you from AI Overviews." That sentence is now wrong on two counts, verified 2026-08-22: (a) Google documents a Search Console "Search generative AI control" that excludes a property from AI Overviews, AI Mode and generative AI features in Discover while explicitly not affecting the rest of Search (https://support.google.com/webmasters/answer/16908024). Rechecked 2026-09-20: Google now states worldwide rollout as of 2026-08-31. (b) Microsoft has documented exactly that split since September 2023 via NOARCHIVE, and states content with NOCACHE or NOARCHIVE "will still appear in our search results". Rewrite that item to point here rather than making the negative claim. 2. /blog/why-law-firm-not-showing-up-chatgpt — cause 4 gains ONE line pointing here, framed as "robots.txt is not the only way your own site turns AI answers away". Do not expand cause 4 beyond one line; that post's six-cause shape is settled. 3. /methodology — the inference-labelling paragraph should link here as the primary-source backing for why Briefly labels its own causal statements as inference. DELIBERATE NON-PROPOSAL. This piece obviously suggests a fourth free tool: a snippet-directive reader that fetches a page and reports its robots meta tags and X-Robots-Tag header. Do not build it. PRODUCT_BETS_AND_OPERATING_STRATEGY.md gates additional free tools behind a measured signup-and-activation path on the existing diagnostic (Bet 4: >=60% check completion, 8-15% to signup, >=50% of those to a first scheduled scan). That gate is unmet. The manual two-minute check in the body is the substitute, and it costs us nothing to give away. -->

Update — September 20, 2026: Removed the outdated limited-rollout description of Google's Search generative AI control. Google now states worldwide rollout on August 31. Rechecked the control and Generative AI performance report documentation; clarified that missing report data is not zero visibility. Other source-verification dates below remain unchanged.

Short answer, from Google's own specification: nosnippet does, and max-snippet limits it. Google documents that nosnippet "will also prevent the content from being used as a direct input for AI Overviews and AI Mode," and that max-snippet "will also limit how much of the content may be used as a direct input for AI Overviews and AI Mode" (robots meta tag specification, page states last updated 2026-03-24; read 2026-08-22).

Two lines of HTML. No crawler is blocked. robots.txt reads as wide open. The free crawler check reports every bot as allowed — correctly, because access is not the problem. The problem happens after the fetch.

That is the boundary this page keeps. The crawler check owns access: which bot can reach you, under which user agent, and what your robots.txt says about it. This page owns everything after the fetch — what each operator documents about how it selects, uses and attributes what it already has, and which publisher controls act at that layer.


Where the directive comes from, and why nobody notices it

nosnippet and max-snippet are old SEO instruments. Publishers set them to control how much of a page Google may reproduce on the results page. News sites set them for licensing reasons. Plugins expose them as a checkbox.

Yoast SEO — one of the most widely installed SEO plugins on WordPress — writes a robots meta tag on every public page by default, in this shape (functional specification, read 2026-08-22):

<meta name="robots" content="{{values}}, max-snippet:-1, max-image-preview:large, max-video-preview:-1" />

max-snippet:-1 is the permissive value — Google's spec defines -1 as "Google will choose the snippet length that it believes is most effective." The default is fine. The default is not the problem.

The problem is the per-page No snippet checkbox on the Advanced tab. Yoast's own help page describes it in exactly one register (meta robots advanced settings, read 2026-08-22):

"If you select No snippet, you prevent the search engines from showing a snippet of this page in the search results. It also prevents search engines from caching the page."

Search results. That is the entire documented consequence, and in 2019 it was the entire actual consequence. The plugin UI says nothing about AI Overviews or AI Mode, because when the setting was written there were none. A vendor who ticked that box on a firm's practice-area pages — to stop Google pulling an awkward sentence into a snippet, to control a licensing exposure, to satisfy a partner who did not like what was showing — made a decision about the 2019 SERP that Google's current specification also applies to its AI surfaces.

How common is this on law firm sites? We do not know. Briefly has not measured it, so this page prints no number. What we can say is what the documents say and how to check your own pages in about two minutes.

Check your own pages

No tool, no signup. In a browser, open a page on your site, view source, and search the HTML for max-snippet and nosnippet. Or from a terminal:

curl -sI https://yourfirm.example/practice-areas/ | grep -i x-robots-tag
curl -s  https://yourfirm.example/practice-areas/ | grep -io '<meta[^>]*robots[^>]*>'

Two places to look, because the directive can arrive either way. The meta tag lives in the HTML; the X-Robots-Tag HTTP header does the same job from the server and never appears in the page source. Check your homepage, your main practice-area pages, your attorney bios and your city pages — the settings are per-page, so one bad template is enough.

What you are looking for:

  • nosnippet anywhere in a robots meta tag or X-Robots-Tag header.
  • max-snippet:0 — Google's spec defines 0 as "No snippet is to be shown. Equivalent to nosnippet."
  • A small positive number, like max-snippet:20. This does not remove you; it caps what can be used.
  • noindex, which removes the page from Search entirely and therefore from the AI surfaces built on Search.

One trap worth knowing: Google states that "in the case of conflicting robots rules, the more restrictive rule applies. For example, if a page has both max-snippet:50 and nosnippet rules, the nosnippet rule will apply." A permissive site-wide default does not rescue a restrictive per-page tag.

And if you remove a directive, nothing happens immediately. Google's troubleshooting guidance for exactly this situation says to "allow time for Google to recrawl and process the change in preview controls," noting that "crawling can take anywhere from several days to several months" (AI features and your website, read 2026-08-22).


The eligibility chain, in Google's words

The reason a snippet directive reaches AI answers at all is that Google routes AI eligibility through snippet eligibility. Three documents, one chain:

  1. AI features are Search. "AI is built into Search and integral to how Search functions, which is why robots.txt directives for Googlebot is the control for site owners to manage access to how their sites are crawled for Search. To limit the information shown from your pages in Search, use nosnippet, data-nosnippet, max-snippet, or noindex controls." (AI features and your website)

  2. Eligibility requires a snippet. "To be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements. There are no additional technical requirements." (AI features and your website, which states last updated 2025-12-10.)

    Note that date, because a second Google page has since gone further. The generative AI optimization guide, last updated 2026-07-10, repeats the first sentence almost verbatim for generative AI features generally — and then names a requirement the older page's "no additional technical requirements" does not: "In addition to the technical requirements for Search, a site must be included in Search generative AI features in Search Console to be eligible for display in generative AI features on Google Search." That is Google's own documentation folding the Search Console control described below into the eligibility gate itself. Include is the default for every property, so this only bites if somebody changed it — which is exactly the reason to go and look.

  3. The directives reach the AI surfaces by name. The robots meta specification lists AI Overviews and AI Mode inside the scope of both nosnippet and max-snippet, and adds the direct-input language quoted at the top of this page.

Nothing in that chain is inference. It is Google's own documents, quoted, and they agree on the chain — where the newer page differs from the older, it adds a gate rather than loosening one.

One real carve-out

max-snippet has a documented exception, and it is the only place in Google's current documentation where structured data changes what an AI surface may use:

"However, this limit does not apply in cases where a publisher has separately granted permission for use of content. For instance, if the publisher supplies content in the form of in-page structured data or has a license agreement with Google, this setting does not interrupt those more specific permitted uses."

Read that precisely. It says a max-snippet cap does not restrict content the publisher supplied through structured data. It does not say structured data is an input to which firm gets named. Those are different claims, and the second one is not documented anywhere we could find.


The control most firms have never heard of

Between the crawler check publishing on 17 August 2026 and this page publishing five days later, one of our own statements went stale. That page said Google's documentation described no control that keeps ordinary search snippets while removing you from AI Overviews. It does.

Search Console's Search generative AI control separates AI-feature inclusion from ordinary Search. Google now states worldwide rollout on August 31, 2026. Check Settings → Search generative AI, including the effective setting inherited from a parent property. Inclusion is the default; an inherited exclusion still needs attention. This is separate from the Google-Extended training control. Changes require propagation time. Google's control documentation, rechecked September 20, 2026.

The Generative AI performance report, also rechecked September 20, reports impressions in AI Overviews and AI Mode. Its rollout notice now says worldwide, but its troubleshooting section still mentions access and insufficient impressions. Record a missing report as unavailable, not zero. For a page-level workflow, use the law-firm Google AI checklist; for attribution beyond impressions, see what intake and analytics can see.

While we are correcting ourselves: Google-Extended is not an access control

Google-Extended appears in a robots.txt file, which makes it look like an access rule. It is not one. Google's crawler documentation is unusually direct about this: "Google-Extended doesn't have a separate HTTP request user agent string. Crawling is done with existing Google user agent strings; the robots.txt user-agent token is used in a control capacity."

It is a use control wearing robots.txt clothing — which is why it belongs on this page rather than the crawler one. What it governs is training and "grounding (providing content from the Google Search index to the model at prompt time to improve factuality and relevancy) in Gemini Apps and Grounding with Google Search on Vertex AI." What it does not govern is stated in the same entry: it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search."

So the practical shape is: Google-Extended can remove you from a grounded Gemini Apps answer and leave AI Overviews untouched. The crawler registry carries the access half of this; the distinction between the two halves is the whole reason both pages exist.


The operator registry

One row per operator. The question each column asks is not "what works" — nobody outside these companies can answer that — but the narrower, checkable one: what does this operator publish, at a URL you can open right now?

Every row links to the operator's own documentation and carries the date we read it. That date is the only freshness claim this page makes.

<a id="operator-registry"></a>

Operator / surfaceDocumented control that acts after the fetchDocumented basis for choosing sourcesPrimary doc (read 2026-08-22 unless noted)
Google Search — AI Overviews, AI ModeYes, several: nosnippet, max-snippet, data-nosnippet, noindex, plus the Search Console Search generative AI control (worldwide rollout stated August 31, 2026). Google-Extended separately governs Gemini Apps grounding and training.Partial. States AI features are "rooted in our core Search ranking and quality systems," describes retrieval-augmented generation and query fan-out. Publishes no criterion for which business is named.Robots meta spec · AI features · Generative AI guide · Search Console control
OpenAI — ChatGPT searchNone documented. The crawler documentation describes robots.txt opt-outs only, and notes "it can take ~24 hours from a site's robots.txt update for our systems to adjust."None documented. The crawler page states no selection or ranking criteria. Its one content-relevance statement is about ads (see below).OpenAI Crawlers
PerplexityNone documented. robots.txt, published IP ranges, and WAF allowlist guidance only.None documented on the crawler page.Perplexity Crawlers
Anthropic — Claude searchNone documented. States its bots "respect 'do not crawl' signals by honoring industry standard directives in robots.txt," plus a non-standard Crawl-delay.None documented. Claude-SearchBot "analyzes online content specifically to enhance the relevance and accuracy of search responses" — a purpose, not a criterion.Anthropic crawler article (article dated April 7, 2026)
Microsoft — Copilot, AI answers in BingYes, and it is the oldest of them. NOARCHIVE content "will not be included in Bing Chat answers, not be linked to in the answers." NOCACHE content "may be included in Bing Chat answers. We will only display URL/Snippet/Title in the answer."None documented as a ranking basis. Bing's AI Performance reporting explicitly disclaims it: its cited-pages metric "does not indicate ranking, authority, or the role of any page within an individual answer."Bing Chat controls (Sept 22, 2023) · AI Performance (Feb 2026)

Verified on 2026-08-22 against the linked documents; Google Search Console control availability rechecked 2026-09-20. Two caveats on the Microsoft row: the 2023 post says "Bing Chat," the product now branded Microsoft Copilot, and Microsoft has not reworded the post; and Google lists both noarchive and nocache under a section headed "Reference of historical and other unused rules," which states "The following rules aren't used by Google Search and are ignored" — so those tags mean one thing to Bing and nothing at all to Google.

The directive table

This is the part worth bookmarking, and the part worth copying into your vendor's ticket.

<a id="snippet-directives"></a>

DirectiveWhat the operator documents for ordinary resultsWhat the operator documents for AI surfaces
nosnippet (Google)"Do not show a text snippet or video preview in the search results for this page."Applies to "all forms of search results (at Google: web search, Google Images, Discover, AI Overviews, AI Mode)" and "will also prevent the content from being used as a direct input for AI Overviews and AI Mode."
max-snippet:[n] (Google)"Use a maximum of [number] characters as a textual snippet for this search result." 0 is "Equivalent to nosnippet"; -1 lets Google choose.Applies to the same list including AI Overviews and AI Mode, and "will also limit how much of the content may be used as a direct input for AI Overviews and AI Mode." Does not apply to content the publisher supplied via in-page structured data.
data-nosnippet (Google)An HTML attribute on span, div and section elements designating "textual parts of an HTML page not to be used as a snippet."Google lists it among the controls that "limit the information shown from your pages in Search," and Search includes the AI surfaces. Note: "structured data remains usable for search results when declared within a data-nosnippet element."
noindex (Google)"Do not show this page, media, or resource in search results."Removes the page from Search, and AI eligibility requires being indexed — so it removes AI eligibility with everything else. The blunt instrument.
Search generative AI control (Google, Search Console)No effect. "This control isn't used as a ranking or inclusion signal affecting other parts of Search."Excludes the property from AI Overviews, AI Mode and generative AI features in Discover. Google now states worldwide rollout on August 31, 2026 (rechecked September 20).
Google-Extended (Google, robots.txt token)No effect. "Does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search."Governs training of Gemini models and grounding in Gemini Apps and Vertex AI. Not a control for AI Overviews.
NOARCHIVE (Microsoft)Content "will still appear in our search results.""Will not be included in Bing Chat answers, not be linked to in the answers."
NOCACHE (Microsoft)Content "will still appear in our search results.""May be included in Bing Chat answers. We will only display URL/Snippet/Title in the answer." If a page carries both tags, "we will treat it as NOCACHE."

Quotations above retain their original 2026-08-22 source-check date; the Search Console control availability statement was rechecked on 2026-09-20. Sources are listed at the foot of this page. Reuse the table if it is useful — please keep this line with it:

Source: Briefly, "Does max-snippet stop AI Overviews? What Google actually documents," https://www.askbriefly.ai/blog/ai-assistant-source-documentation


The honest negative

Here is the finding that matters most, and it is a negative one.

We opened every URL listed in the sources block on 22 August 2026. At those URLs, on that date, no operator documents a mechanism for choosing which local professional-services provider to name in an AI answer. Not Google, not OpenAI, not Perplexity, not Anthropic, not Microsoft. Not for law firms, not for accountants, not for contractors.

That is a statement about documents we read on a date, not a claim about what these systems do internally. The mechanisms exist; they are simply not published. Three URLs we could not read at all are named in the sources block, and nothing on this page characterises their contents.

What the operators do publish stops short in a consistent way:

  • Google describes the architecture — retrieval-augmented generation, query fan-out, "rooted in our core Search ranking and quality systems" — and the eligibility gates quoted above. It does not describe a selection criterion.
  • Microsoft publishes citation counts and then disclaims the reading you would want from them: its cited-pages metric "does not indicate ranking, authority, or the role of any page within an individual answer," and its page-level counts reflect "how often pages are cited, not page importance, ranking, or placement."
  • OpenAI, Perplexity and Anthropic publish crawler purposes and robots.txt behaviour, and stop.

Which means the following, including about us

Every "here is how to get recommended by ChatGPT" claim in this category is inference. Ours included. When Briefly observes that firms with complete directory profiles appear more often in the answers we collect, that is a correlation in our own dataset, labelled as such on our methodology page — not a documented mechanism, and not something any operator has confirmed. The same discipline that makes a vendor's promise checkable makes ours checkable: the promise ladder exists precisely because "we measured this" and "we know why this happens" are different sentences with different evidentiary weight.

Google says a version of this itself, and it is worth quoting against our own interest:

"Be wary of third-party tools that promise ranking success or claim to use 'internal' Google metrics. No third-party tool has access to our internal ranking or AI systems."

Briefly is a third-party tool. We have no access to any operator's internal systems. What we have is a record of what assistants actually answered, on stated dates, to stated questions.

The one documented content-relevance mechanism, and it is the paid lane

There is exactly one place in all of this documentation where an operator says it reads your page content to decide when to show you. It is OpenAI's ads crawler:

"OAI-AdsBot is used to validate the safety of web pages submitted as ads on ChatGPT. When you submit an ad, OpenAI may visit the landing page to ensure it complies with our policies. We may also use content from the landing page to determine when it's most relevant to show the ad to users."

Documented content-to-relevance mechanism: advertising. Documented content-to-relevance mechanism for the organic shortlist: none we could find, on any operator's site, on 22 August 2026. Make of that what you will; we are not going to make more of it than the sentence supports.


Structured data: what is documented, and what isn't

Firms get sold LegalService and Attorney JSON-LD as an AI-visibility measure. Here is what the documentation supports.

Google's structured data gallery documents which types are supported and what they buy you: eligibility for rich results in Google Search. The LocalBusiness page describes a knowledge panel and business carousels in Search and Maps. Read on 2026-08-22, that page contains no mention of AI Overviews, AI Mode, or ranking of any kind.

Google's generative AI guide addresses the question head-on, in a section it titles "Mythbusting":

"Overfocusing on structured data: Structured data isn't required for generative AI search, and there's no special schema.org markup you need to add. However, it's a good idea to continue using it as part of your overall SEO strategy, as it helps with being eligible for rich results on Google Search."

And on the file formats that get sold alongside it:

"You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn't use them."

So: markup is worth having for rich results, and the max-snippet carve-out above means structured data survives a snippet cap. None of that makes it a documented input to being named. The guardrail we state on our own entity check is unchanged and stated the same way here: no chat assistant publicly documents reading JSON-LD. We say that and stop.

Google Business Profile: what is documented, and what isn't

Google's local ranking help page is specific about ordinary local results: "Local results are mainly based on relevance, distance, and popularity," with a blunt note that "there's no way to request or pay for a better local ranking on Google." Read on 2026-08-22, that page says nothing about AI Overviews or AI Mode.

Google's generative AI guide does connect the two, and this is the strongest documented statement on the subject we found:

"Where appropriate, generative AI responses can include product listings, product information, and information about local businesses. Using products like Merchant Center […] and Google Business Profiles can help your products and services to be visible in both AI responses and other Google Search results."

Note what it is and is not. It is Google saying a maintained Business Profile helps visibility in AI responses. It is not a documented ranking input, a weight, or a criterion. Microsoft's equivalent, in the February 2026 AI Performance post, is the same shape: businesses "can register with Bing Places for Business to help ensure that key details such as address, hours, and contact information remain current and eligible for inclusion in AI-generated responses."

Two operators saying an accurate listing helps eligibility. Neither saying how a name is chosen. That is the whole documented state of the art, and any vendor deck that goes further than it is going further than the record.


The skeptic's position, given its due

A fair objection: this is old SEO with a new invoice attached.

For Google's surfaces, Google agrees with the skeptic:

"From Google Search's perspective, optimizing for generative AI search is optimizing for the search experience, and thus still SEO."

The same guide tells site owners they can ignore chunking content, writing specially for AI systems, llms.txt files, and "seeking inauthentic 'mentions'" — most of the tactic list in circulation. If your vendor's AI-visibility plan is a repackaged technical SEO audit, that is not automatically a scam; on Google's own account it is close to the correct plan, and it should be priced like the technical SEO audit it is.

Where the skeptic's position doesn't hold is the specific mechanism this page is about. A nosnippet tag is not an SEO opportunity, it is a live directive with a documented AI consequence that your plugin's UI does not mention. And the assistants outside Google's index — ChatGPT, Perplexity, Claude — are not Google Search, do not run on Google's ranking systems, and publish nothing about how they choose. "It's still SEO" is Google's answer for Google's surfaces. It is not an answer for the rest.

What to do with this

Four things, in order, none of which requires buying anything:

  1. Read your own robots directives on your homepage, practice-area pages, bios and city pages — meta tag and X-Robots-Tag header both. Two minutes, commands above.
  2. If you find nosnippet or a small max-snippet, find out who set it and why. Sometimes there is a real reason. A licensing constraint or a partner's decision about reproduced text is a legitimate trade-off — it is only a mistake when nobody knew they were making it.
  3. Check your Search Console Settings for a Search generative AI control entry. If it exists, confirm it says include, and check whether your property is inheriting from a parent someone else configured.
  4. Ask your vendor one question in writing: "Which of our pages carry nosnippet or a max-snippet limit, and who set them?" Server-side and CDN-injected headers are theirs to answer for; you cannot see all of them from a browser.

Then check the layer this page deliberately does not cover — whether the crawlers can reach you at all — with the free AI crawler check, and the wider set of reasons a firm goes unnamed in the six-cause diagnostic.

FAQ

Does max-snippet stop AI Overviews? max-snippet:0 does — Google's specification defines 0 as equivalent to nosnippet, and nosnippet "will also prevent the content from being used as a direct input for AI Overviews and AI Mode." A positive value like max-snippet:20 does not remove you; Google documents that it "will also limit how much of the content may be used as a direct input." Google publishes no threshold at which a limit becomes an effective exclusion, so we will not invent one.

Is there a way to leave AI Overviews and stay in normal search results? Google documents one: the Search generative AI control in Search Console, which excludes a property from AI Overviews, AI Mode and generative AI features in Discover, and which Google states "isn't used as a ranking or inclusion signal affecting other parts of Search." Google now states worldwide rollout on August 31, 2026; this availability claim was rechecked September 20. Microsoft has documented the equivalent since 2023 via NOARCHIVE, stating that content carrying it "will still appear in our search results." The snippet directives are not this — they affect ordinary results too.

Do AI assistants read my schema markup to decide whether to recommend me? No operator documents doing so. Google's generative AI guide states plainly that "structured data isn't required for generative AI search, and there's no special schema.org markup you need to add," and recommends it for rich-result eligibility instead. We checked OpenAI's, Perplexity's and Anthropic's published documentation on 22 August 2026 and found no mention of reading JSON-LD at all. Markup is worth having; it is not a documented recommendation input.

Why won't you just tell me how to get recommended? Because nobody has published the mechanism, and a page that invented one would be worth less than this one. We checked every URL in the sources block on 22 August 2026 and found no operator documentation of how a local professional-services provider is chosen for an AI answer. Anyone selling that answer — including us, if we ever do — is inferring. The honest version is: fix the documented blockers, then measure what the assistants actually say, repeatedly, and read the trend.

John Rice builds and operates the scan engine behind Briefly, which runs client-style lawyer-hiring questions across ChatGPT, Gemini, Perplexity, Google AI Overviews and Google AI Mode on a recurring schedule and stores every answer it collects. He is not a lawyer; he measures what AI assistants say, with receipts. Every quotation on this page comes from an operator's own documentation, read on the date stated. Methodology · About John

Briefly measures what AI assistants say. It does not rank, rate, or endorse attorneys, and nothing here is legal advice. Being named by an assistant says something about coverage and sources, not about quality of representation.

Sources

The original source review was 2026-08-22. The Search Console control and Generative AI performance report were rechecked on 2026-09-20. Other dates below describe the original review, not a new verification of every source.

Google

OpenAI, Perplexity, Anthropic

Microsoft

Plugin behaviour

Checked and unread — nothing on this page characterises their contents

  • https://help.openai.com/en/articles/12627856-publishers-and-developers-faq — HTTP 403 to automated retrieval on 2026-08-22.
  • https://www.perplexity.ai/help-center/en/articles/20260806-understanding-source-labels — HTTP 403 to automated retrieval on 2026-08-22.
  • https://www.bing.com/webmasters/help/which-robots-metatags-does-bing-support-5198d240 — renders client-side; returned no readable text to automated retrieval on 2026-08-22.

Related: AI crawler check · Local entity check · Why your firm isn't showing up in ChatGPT · What a vendor can promise · Methodology