If you searched this phrase, you are probably trying to answer one specific question: after AI-assisted drafting, what actually stops a fabricated citation from reaching a judge? The honest answer is that the market splits into a few distinct categories, and they are not interchangeable. This guide walks through each category, what it catches, what it misses, and roughly what it costs, so you can pick based on your actual risk rather than marketing copy.
Existence-only checkers
These confirm a citation exists at the reporter/volume/page cited, typically against a public database like CourtListener. This catches the most obvious failure mode — a case that was invented outright, the exact problem in Mata v. Avianca — but does not catch a real case with a fabricated quotation or a real case cited for a proposition it does not support. Products in this category are fast and inexpensive, and the honest ones are upfront about being existence-focused rather than claiming full accuracy verification.
Full AI research assistants with citation features
Products like Thomson Reuters' Westlaw AI-Assisted Research and Lexis+ AI bundle citation handling into a much larger AI legal-research product, priced accordingly (often well into the hundreds of dollars per user per month, frequently bundled with an existing Westlaw or Lexis subscription). The Stanford RegLab/HAI follow-up benchmark, "Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools," measured hallucination rates in this category between roughly 17% and 34% depending on the specific product and query type — meaningfully lower than general-purpose models, but far from zero. Do not assume "the big platform" solves the problem by default.
Word-native fact-checking add-ins
Clearbrief is the best-known example: a Word add-in that checks citations and links facts to the record, priced at a premium (reportedly in the hundreds of dollars per user per month at list pricing) aimed at mid-size and large firms rather than solo practitioners. These tools integrate well into an existing drafting workflow but carry a cost structure that can be difficult to justify for a solo or small-firm budget.
Three-layer deterministic checkers
This is the category we built Citation Safe into: existence, quote-match, and proposition-support, run as three separate deterministic checks rather than one blended AI judgment, with the false-verify rate for each layer published and updated on a public page. If you only remember one filter for evaluating any tool in this list — including us — ask it what its published error rate is. Most, today, do not have an answer.
Case examples showing why the distinction matters
In Park v. Kim, 91 F.4th 610 (2d Cir. 2024), the fabricated citation would have been caught by even a basic existence check — the case simply did not exist. But per the Charlotin hallucination-cases database, a meaningful share of documented incidents involve real cases with fabricated quotations or misapplied holdings, which an existence-only tool would not flag. Knowing which failure mode you are most exposed to should drive which tool category you choose.
Expert perspective
Researchers behind the Stanford RegLab benchmark have specifically cautioned against treating any single hallucination-rate number as a permanent grade: rates shift with model updates, with the specific legal question type, and with how aggressively a tool is asked to synthesize versus retrieve. The practical implication for buyers is to re-test periodically rather than relying on a vendor's rate from a year-old benchmark.
A buying checklist
- Ask for the published methodology behind any accuracy claim, not just a headline percentage.
- Confirm whether the tool checks existence only, or existence plus quote-match plus proposition-support.
- Check pricing against your actual filing volume — per-seat enterprise pricing rarely fits solo or small-firm budgets.
- Ask how often the underlying case-law database is refreshed.
- Test it yourself on a document with a known planted error before relying on it for a live filing.
A common question
Is the most expensive option automatically the most accurate?
Not necessarily, and price is a poor proxy for accuracy in this category specifically because most vendors, expensive or not, do not publish a methodology-disclosed error rate at all. The Stanford RegLab/HAI study found meaningful hallucination rates even among premium, well-established legal research platforms. Ask for a published number before assuming price signals accuracy.
Related reading
- Existence, Quote, and Proposition Checks: What Each One Actually Catches
- Which Legal AI Tools Hallucinate Least?
- How to Check for AI Hallucinations in Legal Briefs
- The Solo Practitioner's AI Tool Buying Checklist
- How CourtListener Works for Lawyers
Check a brief before you file it →
A note on emerging entrants
The legal AI verification market is still young and evolving quickly, with new entrants appearing regularly across all four categories described above. When evaluating a new or unfamiliar tool, apply the same evaluation framework regardless of how new the vendor is: ask for a published, methodology-disclosed accuracy rate, confirm exactly which of the three verification layers (existence, quote-match, proposition-support) it actually runs, and test it yourself on a document with a known planted error before trusting it with a live filing.
Final takeaway
There is no single "best" tool across every category; the right choice depends on your specific practice's filing volume, budget, and which failure mode (existence, quote accuracy, or proposition support) you are most exposed to. Understanding the category distinctions in this guide is the actual prerequisite to making that choice well, more so than any single vendor's marketing claims.
Re-evaluate your tool choice at least annually; both accuracy benchmarks and pricing across this market are moving quickly enough that a comparison from even a year ago may no longer reflect current reality.
Whichever tool you land on, the deciding factor should be published accuracy methodology, not brand recognition or price alone.
Bookmark this comparison and revisit it before your next tool purchasing decision, since the categories described here will remain useful even as specific vendors and pricing shift.
Every category described here solves a real problem; knowing which one you actually have is the hard part, and now you do.
Choose deliberately, verify independently, and revisit the decision periodically as this market keeps evolving.