Blog

How to vet an AI-visibility vendor, including us

John Rice16 min read
<!-- PUBLISHER: LINK OBLIGATIONS THAT LIVE OUTSIDE THIS FILE (design/content/CONTENT_PLAN.md, B6). This file is the only one the author touched. All four items below are somebody else's file: 1. content/blog/track-chatgpt-recommendations-law-firm.md, line ~56 — the paragraph beginning "If you're evaluating any vendor, ours included, ask two questions…". CONTENT_PLAN B6 requires this be cut to a sentence plus a link when B6 ships, so the two pages do not compete. Those two questions are tests 2 and 4 here. Suggested replacement: "If you're evaluating any vendor, ours included, there are seven questions worth asking; we published all seven, with our own answers and our own failures, in [how to vet an AI-visibility vendor](/blog/vet-ai-visibility-vendor-law-firm)." 2. content/blog/what-ai-visibility-vendors-can-promise-law-firms.md — B2 owes this page exactly one inbound link (its own publisher comment already records the obligation). This page links to B2 exactly once, at test 6. Keep it one each way. 3. /methodology and /about — CONTENT_PLAN B6 requires inbound links from both. 4. Resolved 2026-08-25: B2's reuse-attribution line and the scorecard attribution on this page both use the canonical https://www.askbriefly.ai host. Also on ship: the scorecard is the designated linkable asset. Keep #vendor-scorecard and the per-row ids stable, and keep the attribution line directly beneath the table. -->

First, the skeptic is mostly right

The case against this category deserves to be stated by someone inside it, at full strength.

Most AI-visibility offerings are the previous decade's SEO retainer with a new noun on the invoice. The deliverables are the same: directory cleanup, content production, review solicitation, schema markup. The reporting is a dashboard with a number that goes up, and that number is new, proprietary, and not defined anywhere you can read it. The pitch arrives with a screenshot of ChatGPT naming a competitor, which is an emotionally effective artifact and an evidentially worthless one. A managing partner who concludes "this is a scam, next" is making a defensible call on the evidence in front of them.

Here is the part that is not a scam. Whether an assistant names your firm when someone asks who to hire is a fact about the world: you can ask the question, write down the answer, ask again tomorrow, and count. It is a tedious fact to observe well, which is the whole reason a tool can be worth paying for. But it is a fact, and facts can be checked.

So the useful question is narrower than "is this category real". It is whether this vendor's number is a measurement or a rendering. Seven tests settle it, all of them over email, before you take a demo. Our own answers are at the bottom, failures included.


The seven tests

1. What is the formula, and where is it published? {#test-formula}

The failure mode is the blended score: mentions, citations, average position and coverage combined into one figure out of 100, with the weights never stated. Blending is not the problem. Unpublished blending is.

An unpublished formula cannot be reproduced, so it cannot be checked, and it cannot be compared to anything, because nobody else computes it the same way. It can also be revised without notice. So when the score moves you cannot tell whether your firm moved or the weights did. One party in the room knows, and it is not you.

Ask: "Send me the arithmetic for your headline number. Not the inputs, the arithmetic."

A real answer is a public page with a formula on it, specific enough that you could recompute the number by hand from their raw data. A list of ingredients is not a recipe: "mentions, position, citations and coverage" tells you nothing about what a six-point move means.

2. How many runs per question, per assistant? {#test-runs}

The failure mode is a single run sold as a measurement. Ask an assistant the same hiring question twice and you do not reliably get the same firms back, so one run is a coin flip with a decimal point on it.

That instability has been measured. A 2026 audit of retrieval-augmented commercial recommendation found identical reruns of the same prompt produced recommendation sets overlapping only 0.50 to 0.61, and cosmetic rewordings of the same intent dropped overlap to 0.288 (arXiv 2605.27440). Disclosure, since this page is partly about disclosure: unreviewed preprint, authors who sell in this category, measured on consumer buying questions rather than legal ones. We cite it against our own interest, because it argues per-prompt tracking is shakier than tools like ours imply.

Runs also have to be counted per assistant rather than pooled. Only 11% of cited domains appeared across more than one of four platforms tested in Whitehat SEO's 118,000-answer analysis (March 2026; Whitehat sells search-visibility services).

Ask: "How many runs per question, per assistant, and what does your product do below that number?"

A real answer is a specific integer with a specific behaviour attached. Watch for frequency substituted for repetition: "we scan daily" says how often the job runs, not how many times each question was asked, and only one of those stabilises a percentage. Our floor is three, enforced in the free answer log, where a share refuses to render until every pair reaches it.

3. Can I have this as data instead of a screenshot? {#test-screenshots}

The failure mode is the screenshot as evidence. A screenshot establishes that one answer existed once, in one session, on an account that may have been personalised by location and history, possibly after several regenerations that were not screenshotted. It cannot be counted, sampled, sorted or re-run.

This is not an argument against screenshots but about their job. A screenshot attached to a row in a dataset is a caption. A folder of screenshots offered instead of a dataset is a mood.

Ask: "Send me the same finding as a CSV, one row per answer."

A real answer arrives as a file or an export button, one row per answer, carrying the date, assistant, question, outcome and sources cited.

4. Can I read the raw answers behind every number, without asking you? {#test-raw-answers}

The failure mode is a dashboard with nothing underneath it. If the only artifact is the aggregate, you are not auditing a measurement. You are auditing the claim that one happened.

The word doing the work is without. "We can pull that for you" is a different product from "click the number". Evidence you have to request is evidence you cannot check on the day you doubt the number, which is the only day it matters.

Ask: "Show me the click path from a headline number to the full text of one answer behind it."

A real answer is a live click during the demo, not a roadmap item.

5. What is in the denominator, and what happens to a failed scan? {#test-denominator}

The failure mode is quiet, technical, and the most consequential here. Scans fail: timeouts, refusals, errors, empty responses. If a failed scan counts as "your firm was not recommended", every assistant outage renders on your chart as a decline and every recovery as an improvement. Notice which way that bias points: it manufactures a dip before the engagement and a recovery during it. A vendor need not have designed it that way to benefit from it, which is exactly why you ask rather than accuse.

The companion rule is what a zero means. A zero should mean "we measured, and it was zero". If it can also mean "we did not measure", the two are indistinguishable on the chart and one of them is a lie of omission. Unmeasured should render as an em-dash and nothing else.

Ask: "Are failed scans excluded from the denominator? How many failed last month?"

A real answer excludes them and reports the count separately, unprompted. Ours is published on the Case Recommendation Share page.

6. Name the mechanism by which you place my firm in an answer {#test-mechanism}

The failure mode is guaranteed placement. This test asks for a mechanism, and its value is that there is no true answer to give.

Be precise, because the sloppy version of this argument is now factually wrong. Paid placement inside assistants exists. OpenAI runs ads in ChatGPT and states that ads are "clearly labeled as sponsored and visually separated from chatbot responses" and that "advertising will not influence answers generated by ChatGPT" (Search Engine Land, March 2026; OpenAI's own pages return HTTP 403 to automated retrieval, so we quote reachable reporting rather than characterise a document we could not open). Perplexity ran sponsored follow-up questions from November 2024 — labelled "sponsored", with answers "generated by Perplexity, not the brand sponsoring the question" (Search Engine Land) — and then ended the programme in February 2026, phasing the ads out with no plans to bring them back, on the stated view that sponsored placements risk the trust an answer engine depends on (Search Engine Land). Operator-side facts move. Check the date on any claim about them, including these two.

So the accurate sentence is not "there are no ads in ChatGPT". It is that the sponsored slot is labelled, held separate from the answer, and — in ChatGPT today — buyable, while neither operator we can source, OpenAI or Perplexity, describes any paid route into the organic recommendation. We have not verified Google's published position for AI Overviews, AI Mode or Gemini, so we assert nothing about it. That gap is the test rather than a hole in it: a vendor promising placement in the unsponsored shortlist is describing a lever it should be able to name, sold by an operator it should be able to name, and you can check both answers against the operator yourself.

Ask: "Which lever moves the organic answer, and who operates it?"

A real answer declines the premise: we cannot place you, we can work the source layer and measure whether the rate moves. What a vendor may safely promise a firm governed by Rule 7.1 is a separate question, answered in what an AI-visibility vendor can legally promise your law firm.

7. Who produced the evidence in this deck, and what do they sell? {#test-evidence}

The failure mode is an evidence base that belongs to the seller, in two shapes.

The first is vendor-published research quoted as though independent. Research produced by a company selling into the category is not automatically wrong, and it is frequently the only research that exists, because nobody else funded the fieldwork. It just has to be labelled, every time, so you can weigh it.

The second is the community thread: somebody posts that their firm has gone invisible in ChatGPT, helpful replies arrive naming a tool. Threads cannot be audited, and that is as true of genuine threads as of the other kind. So the rule is not "assume astroturf", it is that a thread is not a source, and the only correct response to one is to ask for the underlying document. That includes threads mentioning us.

Ask: "For each statistic in this deck: who ran it, when, on what sample, and do they sell in this category?"

A real answer carries a name, a date, a sample size and a disclosure line. Here are ours, since a rule a vendor exempts itself from is just marketing. Every outside finding we lean on comes from an organisation that sells into this category, labelled where it appears: two in test 2 above, plus InterCore and 5WPR/Haute Lawyer, both of which sell to law firms. The InterCore finding, in the form we always state it: a legal directory was the first source cited in 77.8% of AI answers to lawyer-hiring questions (1,254 of 1,612 valid answers, 540 queries asked three times), July 2026. The rest sits on our statistics hub.


Our own answers, failures included

Our August 22 review recorded passes on 4, 5 and 6, partials on 1, 2 and 3, and a failure on 7. The row-by-row version is in the scorecard below. The September 7 corrections explain what has changed since that review and what remains unproven.

Correction, September 7, 2026: the revenue projection described in our August audit has been removed. That projection combined assumptions into an ROI multiple and projected case revenue. The free-check result now explains the observed findings and their limits instead. The Case Recommendation Share formula remains public, with its classification rules and a worked example. A single free check is still a snapshot, not a stable recommendation rate or a measurement of lost revenue.

We fail test 7 outright. No benchmark study, no customer counts, no case studies. No Briefly-generated statistic appears anywhere on this site, deliberately, because the first data study is still in fieldwork. That is the correct posture, and from a buyer's chair it is still an absence of evidence.

Pricing is now public. Our August audit also flagged the lack of a pricing page. You can now compare Briefly's plans, limits, and trial terms before creating an account. This corrects a buying-information gap; it does not supply the independent outcome evidence we still lack.

The passes are on the methodology page, and our editorial policy is the standing version of this section. If a line here stops being true it should be corrected, not quietly dropped.


The scorecard {#vendor-scorecard}

Send this to three vendors. Keep our column or delete it, and add one per vendor. The point is not that our answers are the right answers. It is that a vendor who will not answer in writing has told you something.

#The question to sendWhat a real answer containsBriefly, as of 2026-08-22
1Send me the arithmetic for your headline number.A public formula you could recompute by hand.Published for Case Recommendation Share. September 7 correction: the assumption-based revenue projection has been removed.
2How many runs per question, per assistant, and what happens below that?An integer, plus the behaviour below it.Three, enforced in code; below it no share renders. The free check itself is one pass, so it is a snapshot.
3Send me the same finding as a CSV, one row per answer.A file, not a deck.Answers export with date, assistant, question, outcome, sources. The question set is not yet published.
4Show me the click path from a number to one raw answer.A live click, not a roadmap item.Every metric opens to the stored answer behind it.
5Are failed scans excluded from the denominator? How many failed last month?Excluded, and counted separately, unprompted.Excluded entirely; failures reported on their own. A zero always means measured-zero.
6Which lever moves the organic answer, and who operates it?A refusal of the premise.No placement promised. No vendor can promise one.
7For each statistic: who ran it, when, on what sample, do they sell in this category?Name, date, sample, disclosure.Third-party statistics are linked and labelled. Our own published measured evidence: —

Em-dash means not measured. It never means zero.

The scorecard records the August 22 review, with the September 7 product correction marked above. The content update is not a new verification of every external study or operator policy cited in that review.

Reuse this scorecard. Copy it into your diligence file or your procurement template, and please include this line: Source: Briefly, "How to vet an AI-visibility vendor, including us," https://www.askbriefly.ai/blog/vet-ai-visibility-vendor-law-firm

How the outside claims here were checked. Every external statistic on this page was fetched and read on August 22, 2026, and anything unreachable is not asserted: openai.com, help.openai.com and perplexity.ai each returned HTTP 403, so those advertising positions are quoted only through named secondary reporting. No competitor is named here for a negative finding, because a "we checked and they don't publish theirs" claim built on a marketing homepage is the same unsubstantiated specificity this page argues against. That is what the scorecard is for: the answer should come from the vendor, in writing, to you.


Before you start a paid trial

Evidence quality is only part of the buying decision. A managing partner, marketing director, or agency also needs to know what work the subscription covers and who will use the findings.

Buying questionAsk the vendor to show
Will this cover our firm or agency roster?Included firms, locations, practice areas, competitors, and assistants, with the limits for the quoted plan.
What happens on a schedule?The actual scan cadence, repeats per question, and what happens when a run fails.
Can we use the findings in our workflow?A completed report, the raw-answer view, available exports, and who can access them.
What will we pay after the trial?Recurring price, billing period, usage limits, trial end date, and cancellation process in writing.
What would make us keep the subscription?A trial checklist agreed by your team: collect a baseline, inspect sources, export findings if required, and assign follow-up work.

For Briefly, start with public pricing and plan details and compare them with your requirements. Judge the trial on whether the workflow produces evidence you can use. A short trial cannot establish that the tool caused more recommendations, inquiries, or signed clients.


FAQ

Is AI visibility a scam? The category contains a great deal of repackaged SEO sold on an undefined number, and a buyer who assumes that by default will be right more often than wrong. Whether an assistant names your firm is still a measurable fact. The seven tests separate vendors measuring it from vendors rendering it, over email, before any demo.

What if a vendor answers all seven well? Then their number is real, which is a different question from whether their work moves it. No honest measurement vendor can answer that one: if your recommendation rate rises next quarter, their work is a plausible cause and not a proven one.

Can I run these tests without buying anything? Yes, and run test 2 on yourself first. Ask your own hiring questions three times each across the assistants your clients use and log every answer. Our free answer log does the arithmetic and enforces the run floor; the manual method is in the tracking guide. An hour of that makes every vendor demo afterwards much shorter.

John Rice builds and operates the scan engine behind Briefly, which runs client-style lawyer-hiring questions across ChatGPT, Gemini, Perplexity, Google AI Overviews, and Google AI Mode on a recurring schedule and stores every answer it collects. He is not a lawyer; he measures what AI assistants say, with receipts. The seven failure modes on this page are the ones visible from inside a scan engine, and the ones Briefly currently fails were read out of our own source and our own published pages rather than recalled. Methodology · About John

Briefly measures what AI assistants say. It does not rank, rate, or endorse attorneys, and nothing here is legal advice.

Sources

Research, with disclosure

Assistant operators' advertising positions. openai.com, help.openai.com and perplexity.ai return HTTP 403 to automated retrieval. Their positions are quoted above only through the named secondary reporting below.

Pricing

  • Gauge, published pricing: Growth $599/month, Enterprise custom (read August 22, 2026): https://withgauge.com

Related on this site

<!-- The companion piece on what a vendor may PROMISE is linked exactly once, at test 6, and is deliberately absent from this Related row: B2 and B6 carry one cross-link each way. -->