How to Track When ChatGPT Recommends Your Law Firm
There are two real ways to find out whether AI assistants recommend your firm: ask them yourself, on a schedule, and keep records — or pay a tool to do it. This guide covers both, starting with the free one.
One promise up front: every statistic below links to its source, and where a stat has limitations, we say so.
Want to see your firm's current answers first? Run the free law-firm AI visibility check, with no signup. Use that snapshot to inspect named firms and cited sources, then follow the method below to repeat your questions and record changes. If you already maintain a baseline, skip to the comparison of manual and paid tracking.
Why tracking AI recommendations is worth a partner's time
Three numbers make the case. Each has caveats, noted plainly.
41.9% of consumers say they would use ChatGPT to research a lawyer for an important legal issue — up from 9% in 2023, according to iLawyerMarketing's 2026 survey of 1,110 U.S. consumers. The same survey found 9.5% would use only AI sources, and that adults 45–60, the bracket that hires lawyers, lead adoption: 76% say they would use AI to research law firms. Caveat: this is stated preference, not observed behavior. But the trend line across four annual editions points one direction.
Monthly AI-assistant-referred sessions grew 9.9x, from 65,249 in November 2024 to 644,478 in May 2026, across the 166 websites in Previsible's AI Traffic Report. That dataset spans multiple industries, not just legal. This is observed behavior, measured in analytics, not a survey.
In one Seer Interactive case study, visitors referred by ChatGPT converted at 15.9%, versus 1.76% for Google organic traffic on the same site (Seer Interactive, 2025). This is a single unnamed client, industry undisclosed, and the AI-referred volume was tiny (about 0.07% of the site's organic traffic). Treat it as a signal, not a law of nature. The mechanism is plausible: someone who clicks through from an AI conversation already did their comparison shopping inside the conversation. They arrive decided.
These studies give a law-firm marketer a reason to investigate the channel. They do not establish your firm's referral volume, conversion rate, or return on tracking. Those require your own records.
Recommendation tracking is the upstream instrument. To measure the people who actually reach the firm, pair it with a structured intake question and GA4’s AI Assistant channel; the free Law Firm AI Referral Intake Kit supplies the question, CRM values, and monthly worksheet.
The manual method: ask the assistants yourself
Start here. It's free, it takes an afternoon, and it will teach you more about this channel than any sales deck. We recommend it even though we sell the alternative.
Step by step
- Write down 10–15 questions a real client would ask. Not marketer phrasing — client phrasing. Real people don't type "personal injury attorney services." They ask messy, situational questions:
- "best car accident lawyer in [your city]"
- "I was rear-ended in [your city] and the insurance adjuster is lowballing me. Should I get a lawyer? Who?"
- "who is a good divorce attorney in [your city] for a contested custody case"
- "got a DUI in [your city] last night, do I need an attorney and who should I call"
- "affordable estate planning lawyer near [your city] for a will and trust" Cover each practice area you care about and each phrasing style: the terse search-style query, the story-with-a-question, and the "should I even hire a lawyer" question that precedes the shortlist. Short on phrasings? Download hiring questions for your metro — free, no signup, with the county, metro-nickname, and state-law wording already written in, as CSV or plain text.
- Open a fresh, logged-out session for each assistant. Your own ChatGPT account has memory. It may know you're a lawyer, and that contaminates the answer. Use a private browser window or a logged-out session so you see something closer to what a stranger sees.
- Run each question on each assistant. At minimum: ChatGPT, Gemini, and Perplexity. Then run the same questions as Google searches and note what AI Overviews and AI Mode say. For many clients, that's the first AI answer they ever see.
- Record four things per answer: which firms were named, in what order, whether yours appeared, and which sources the assistant cited or leaned on. Screenshot everything. Copy the full answer text into a document; answers get regenerated, and you will want the receipt.
- Log it in a spreadsheet — or record the answers in the free answer log. One row per question-per-assistant-per-date. Columns: date, assistant, question, firms named (in order), your position or "absent," sources cited. The log does that structure and the share arithmetic for you, in your browser; nothing you type is sent anywhere.
- Repeat every run at least three times. Not once. This matters more than anything else in this list, for reasons the next section explains.
- Re-run the same set monthly. Keep the prompts and sampling method consistent, record changes, and accumulate several periods before interpreting a movement as a trend.
Do this honestly and you'll have something most of your competitors don't: actual evidence of what AI assistants say about your market.
Either way, hold the three-run floor: no recommendation share for a question/assistant pair you have run once. That floor is a reporting rule, not a guarantee of statistical certainty.
Where the manual method breaks
We built a product because we hit these walls ourselves. In order of severity:
The answers aren't stable. Ask the identical question twice and you can get different firms. SparkToro's research, published January 28, 2026, found AI assistants highly inconsistent when recommending brands and products across repeated runs. A single query proves almost nothing. A firm that appears in three of ten runs is in a very different position than one that appears in ten of ten, and one manual check can't tell them apart. "I asked ChatGPT and we came up" is an anecdote with a sample size of one.
Small wording changes swing the answer. "Best car accident lawyer in Denver" and "who should I hire after a car accident in Denver" can produce different shortlists. You need paraphrase coverage, which multiplies the work.
Each assistant is its own market. ChatGPT, Gemini, Perplexity, Google AI Overviews, and Google AI Mode use different models, different retrieval, different sources. Coverage on one says nothing about the others. Five assistants × 15 questions × 3 repetitions is 225 answers to run and log. Per market. Per month.
Models change under your feet. A model update can reshuffle recommendations overnight, silently. Without a before/after record, you can't even detect that it happened.
There's no history unless you build it. The spreadsheet only shows a trend if someone actually maintains it: every month, without shortcuts, across every cell of that matrix. In our experience, that someone stops existing around week six.
Manual vs. spreadsheet vs. purpose-built: an honest comparison
| Manual spot-check | Spreadsheet system | Purpose-built tool | |
|---|---|---|---|
| Cost | Your time | Your time for each scheduled collection | Subscription plus time to review findings; verify scope and limits |
| Measurement limits | One answer cannot establish a rate | Depends on prompt coverage, repetitions, and consistent collection | The same limits apply; automation alone does not establish validity |
| Assistants covered | Whatever you have patience for | All five, at heavy time cost | All supported assistants, automatically |
| Position + source capture | Ad hoc screenshots | Yes, if logged rigorously | Structured, on every answer |
| History / trend line | No | Yes, while discipline holds | Yes, automatic |
| Competitor tracking | Incidental | Manual re-logging | Built in |
| Honest weakness | Anecdote, not data | Dies of neglect by month two | Costs money; you must verify the vendor actually stores raw answers |
| Right for | First look; building conviction | Single-market firms with a diligent marketer | Multi-market firms and agencies |
The spreadsheet column is not a strawman. A disciplined marketing coordinator can genuinely run it for one market. The failure mode isn't capability — it's persistence.
What a systematic approach requires
Whether you build it or buy it, a tracking system that produces evidence rather than anecdotes needs five properties:
- Repetition. The same question, run multiple times per period, because single answers are unstable. Frequency of appearance is the real metric.
- Multiple assistants. All five that matter for consumer legal queries: ChatGPT, Gemini, Perplexity, Google AI Overviews, Google AI Mode.
- Position and competitors, not just presence. "Named first" and "named third after two competitors" are different outcomes. Every answer that omits you names someone else; that's the data you act on.
- Source capture. Which directories, review profiles, and articles the assistant cited. This is the lever you can actually pull, so a system that discards it is decorative.
- Stored raw answers. Every metric should trace back to a real answer you can open and read. If you can't audit the underlying answer, you're trusting a black box to report on a black box.
This is how we built Briefly's engine: we run each firm's question set on a schedule across ChatGPT, Gemini, Perplexity, Google AI Overviews, and Google AI Mode, store every full answer, and extract which firms were named, in what order, and from which sources. We're a measurement company, not a ratings service. We report what the assistants said, with receipts. Our full prompt templates, model versions, and limitations are public on our methodology page.
The metrics that matter
Once answers are flowing in, resist the "AI visibility score" trap: a single opaque number computed who-knows-how. Track these instead:
- Case Recommendation Share — of the client hiring questions in your markets, the share of answers where the AI recommends your firm. The headline number. Full definition and formula here.
- Recommendation share of voice — your recommendations measured against every competing firm named in the same answers.
- Average recommendation position — when named, are you the lead recommendation or the "also consider"?
- Platform coverage — which of the five assistants recommend you, and which ones a competitor currently owns.
- Source presence — how often the sources behind answers point to your site, profiles, and reviews.
Each of these is computable from a well-kept spreadsheet. That's deliberate. A metric you could audit by hand is a metric you can trust from a vendor.
Reporting these to a client, a partner, or a managing committee? Personalize a report template: a free, blank Markdown report that keeps the denominator, the limits, and the raw answers attached to every number. You can fill in the header — firm, market, practice area, window — in your browser, or take it blank. It generates no measurements; the numbers come from your own log.
The tool landscape, including our competitors
Fair picture of the market as of August 2026:
Horizontal AI-visibility platforms — Otterly, Profound, Peec, AthenaHQ, and Trakkr — track brand mentions across AI assistants for any industry. All are legitimate products, and Profound in particular has published some of the most-cited citation research in the space. On price, as of August 2026: Otterly starts at $29/month (its Lite tier, 15 tracked prompts); AthenaHQ's self-serve plan starts at $295/month; Profound runs from $99/month for a ChatGPT-only starter to $399/month for multi-engine coverage, billed annually, with enterprise custom-priced above that. Prices move fast in this category, so confirm on the vendor's page before budgeting. The shared trade-off for a law firm is generality: you write your own prompts, and the tooling doesn't know that legal hiring questions are local, practice-area-specific, and phrased like emergencies ("got a DUI last night") rather than like shopping.
Briefly — that's us — is built only for law firms. You enter your firm, markets, practice areas, and competitors; we generate realistic client hiring questions per market and run them on a schedule across all five assistants, with every answer stored and every metric traceable to the answers behind it. We won't pretend neutrality here, so weigh our description accordingly. And note what we're not claiming: no customer counts, no case studies, no "firms like yours saw X%." The product is new and we'd rather show you than tell you. The free check runs a real scan for your firm in about 60 seconds, no signup — the fastest way to see whether any of this matters in your market before spending a dollar.
If you're evaluating any vendor, ours included, use the seven-question AI visibility vendor scorecard. It covers raw answers, repetitions, failed scans, formulas, and evidence, with our own answers included.
When does paid tracking make sense?
Consider it when you have several markets or client accounts to check, need a history that survives staff changes, and have someone responsible for acting on findings. Keep the manual method if you can maintain it reliably and the paid workflow does not save meaningful work.
Before buying, write down your required markets, practice areas, assistants, scan frequency, and export needs. Ask the vendor to show one completed scan and explain exactly what the quoted plan includes. For Briefly, compare tracking plans and pricing, then use the trial to check the workflow against that list. A successful trial shows you can collect, inspect, and use the evidence; it does not need to show an increase in recommendations or signed cases in a few days.
What to do with what you find
Tracking is diagnosis. The treatment follows from the sources column of your spreadsheet, and the pattern there is well documented.
A legal directory was the first source cited in 77.8% of AI answers to lawyer-hiring questions (1,254 of 1,612 valid answers, 540 queries asked three times) — InterCore Research, July 2026. (Study link.) InterCore sells AI-visibility services; this is vendor-published research, cited here with that caveat.
Separately, a 2026 audit by 5WPR and Haute Lawyer found that a tight set of roughly seven directories — Chambers, Legal 500, Super Lawyers, Best Lawyers, Martindale, Avvo, and Justia — dominated AI citations across every legal query category tested, with zero law-focused editorial sources appearing in top results (report; announcement). Assistants don't know your firm directly. They know what the sources they retrieve say about your firm.
So the playbook is unglamorous:
- Fix the directories your scans actually surface. Not the directories with the best sales team — the ones appearing in your answers' citations. Complete profiles, consistent name-address-practice-area data, current attorney bios.
- Mind your reviews where the assistants look. Google reviews and Avvo ratings show up in cited sources constantly. Volume, recency, and specifics matter.
- Publish pages that answer hiring questions directly. Assistants cite pages that resolve the query. A clear page on "what a contested custody case costs in Ohio" is citable; a homepage slider is not.
- Re-scan, and attribute honestly. Change one thing, watch the next month's scans. Movement in AI answers is noisy; repetition is what separates a real shift from model weather.
Frequently asked questions
How do I check if ChatGPT recommends my law firm right now?
Open a logged-out ChatGPT session and ask 5–10 questions a real client would ask: "best [practice area] lawyer in [city]," plus situational versions. Note which firms are named and repeat each question a few times, because single answers vary.
Can Google Analytics tell me when AI assistants mention my firm?
No. Analytics only records visits — clients who clicked a link from an AI answer, visible under referrers like chatgpt.com. Most AI recommendations end without any click, and a mention with no link leaves no trace. Analytics measures the aftermath of a fraction of answers; only asking the assistants measures the answers themselves.
How often should a law firm scan AI assistants?
Weekly, at minimum monthly. Model updates can reshuffle recommendations without notice, and infrequent snapshots can't separate trend from noise. Whatever the cadence, run each question multiple times per cycle.
Why does ChatGPT give different answers to the same question?
The models are probabilistic — sampling variation alone changes outputs between runs — and retrieval, model version, session context, and phrasing all shift results further. Independent research has found AI assistants highly inconsistent in brand recommendations across repeated identical prompts. This is why appearance rate across repeated scans, not any single answer, is the meaningful measurement.
Is AI recommendation tracking worth it for a small firm?
Do the free version first: an afternoon of manual scans tells you whether AI assistants are naming competitors in your market. If competitors are being named and you aren't, you have your answer. If your scans come back empty of competitors too, save the money and re-check quarterly.
About the author. John Rice builds and operates the scan engine behind Briefly, which runs client-style lawyer-hiring questions across ChatGPT, Gemini, Perplexity, Google AI Overviews, and Google AI Mode on a recurring schedule and stores every answer it collects. He is not a lawyer; he measures what AI assistants say, with receipts. Metrics in Briefly trace back to stored, readable AI answers.
Briefly measures what AI assistants say. It does not rank, rate, or endorse attorneys.
Sources
- iLawyerMarketing, What Online Sources Do People Use to Research and Find Attorneys in 2026? (n=1,110)
- Previsible, AI Traffic Report (July 2026)
- Seer Interactive, How Traffic from ChatGPT Converts (2025)
- InterCore Research, State of AI Search Visibility for Law Firms 2026 (July 2026)
- 5W Public Relations & Haute Lawyer, 2026 Legal AI Visibility Report
- 5WPR press release
- SparkToro, AIs Are Highly Inconsistent When Recommending Brands or Products (January 28, 2026)