Trust

Methodology

Briefly is a measurement company: we report what AI assistants said in response to lawyer-hiring questions, with the underlying answers stored and readable. This page describes how those measurements are produced. It is written in capability terms — what the system does — because our first public data study has not shipped yet, and we won't quote numbers before it does.

What we ask

For each tracked firm, the engine maintains a set of client-style hiring questions scoped to the firm's markets and practice areas — phrased the way real people ask ("got a DUI in Austin last night, do I need a lawyer and who should I call"), not the way marketers search. Question sets mix terse search-style queries, situational story-with-a-question phrasings, and paraphrase variants, because independent research shows wording alone can change which firms an assistant names.

Where and how often

  • Assistants covered: ChatGPT, Gemini, Perplexity, Google AI Overviews, and Google AI Mode.
  • Repetition: questions run on a recurring schedule, with repeated runs per question, because single answers are unstable. Appearance rate across repeated runs — not any single answer — is the unit of measurement.
  • Clean context: scans approximate a clean-context user rather than a personalized session. Real clients have location and history that shift answers; this is a documented limitation, not a hidden one.

How answers are classified

Every successful answer is classified per firm on a five-state ladder — absent, mentioned, cited, recommended, primary recommendation. Only the last two count as recommendations. The full definitions, formula, and a worked example live on the Case Recommendation Share page, published so anyone can compute the metric themselves and check ours.

Failed scans — timeouts, refusals, errors, empty responses — are excluded from denominators entirely. An assistant outage says nothing about a firm, and counting it as "not recommended" would manufacture volatility.

What we store

Every full answer text, with its date, assistant, question, the firms named and their order, and the sources the assistant cited or leaned on. Every metric we display traces back to stored answers a customer can open and read. If a number can't be audited against raw answers, we don't ship it.

Known limitations

  • Assistant outputs are probabilistic; identical prompts produce varying answers. Trends over repeated runs are meaningful, single snapshots are not.
  • Model updates can reshuffle recommendations overnight, without notice.
  • Question realism is a judgment call. We document our question-design approach so disagreements are about the rules, not hidden behavior.
  • Measured answers approximate a clean-context user; no real client is exactly that user.

The forthcoming data study

Our first public data study — a recurring index built from scheduled scans of client-style hiring questions — is in fieldwork. When it publishes, it will carry its full prompt templates, model versions, date windows, sample sizes, and limitations alongside the findings. Until then, no Briefly-generated statistics appear anywhere on this site; our published pages cite third-party research only, each stat to its primary source.

Questions and corrections

Methodology questions, replication requests, and corrections: support@briefly.app. See also the editorial policy and about Briefly.

Briefly measures what AI assistants say. It does not rank, rate, or endorse attorneys, and its measurements are of AI outputs — not of lawyer quality.