Citation verification is often talked about as a single step. It is actually three distinct checks, each catching a different failure mode, and a verification process that only runs one or two of them will miss real, documented categories of AI-hallucination error.
Existence check
This confirms the cited case exists at the specific reporter, volume, and page given. It is the fastest check to run and catches the most obvious failure mode — a case invented outright, as in Mata v. Avianca, No. 22-cv-1461 (S.D.N.Y. 2023), where several cited cases could not be located in any reporter. CourtListener's free citation lookup is the standard free tool for this check.
Quote-match check
This confirms that any language attributed to a case in quotation marks actually appears in the opinion, word for word. A case can pass the existence check and still carry a fabricated quotation — this is a distinct failure mode documented in multiple cases in the Charlotin AI Hallucination Cases database, where the citation was to a real case but the quoted language was invented.
Proposition-support check
This confirms the case actually stands for the specific legal proposition it is cited for, as opposed to being real, accurately quoted, but simply irrelevant or contrary to the point being made (for example, citing a dissent, a vacated holding, or a portion of an opinion that does not address the point at all). This is the hardest check to automate fully and the one most often skipped under deadline pressure.
Why running only one or two is not enough
Existence-only tools, common in the market, will pass a real case with a fabricated quotation or a real case cited for the wrong proposition. Quote-matching without a proposition check can still let through a case that says the quoted words but doesn't actually support your argument's premise. All three checks, run independently, are needed to catch the full range of documented failure modes.
Expert perspective
This three-layer framework is exactly why we built Citation Safe around three separate deterministic checks rather than one blended AI judgment, publishing the false-verify rate for each layer independently rather than a single combined accuracy number that could obscure which failure mode is actually being caught.
A quick reference
- Existence check: does the case exist at this reporter, volume, and page?
- Quote-match check: does this exact quoted language appear in the opinion?
- Proposition-support check: does the case actually stand for what it's cited for?
A common question
Which check catches the most documented errors?
Existence checks catch the largest single category, since outright fabrication is the most common failure mode in the Charlotin database. But quote-match and proposition-support failures are common enough, and dangerous enough on their own, that skipping them leaves real exposure.
Related reading
- How to Check for AI Hallucinations in Legal Briefs
- The Hidden Risk of AI-Generated Quotations in Legal Writing
- AI Legal Citation Verification Tools, Compared
- How CourtListener Works for Lawyers
- The AI Legal Brief Due-Diligence Checklist
Check a brief before you file it →
A worked example across all three checks
Take a hypothetical citation to "Rivera v. State Farm Mutual, 512 F.3d 88 (3d Cir. 2011)" offered for the proposition that an insurer's bad-faith duty extends to pre-litigation settlement conduct. An existence check confirms whether this case is real at all. If it passes, a quote-match check confirms whether any quoted language attributed to the case actually appears at the cited page. If that passes too, a proposition-support check confirms the case is actually about pre-litigation bad faith and not, for instance, a procedural ruling on removal jurisdiction that happens to involve the same insurer. Each layer catches a distinct, independent failure, and a citation can fail at any one of the three stages while passing the others.
Why proposition-support is the hardest to automate fully
Existence and quote-match are largely mechanical: does a string appear in a database, does a text fragment appear in a document. Proposition-support requires actual legal judgment about what a case holds and whether that holding extends to your specific factual and legal context, which is inherently harder to fully automate and is why even sophisticated verification tools generally flag proposition-support results for human review rather than presenting them with the same binary confidence as existence and quote-match results.
Final takeaway
Memorize the three-layer framework, not just the general instruction to "verify citations." Knowing specifically which of the three checks you have and haven't run for a given citation is what separates a thorough verification process from a false sense of having checked something that was only partially reviewed.
How this maps to tool design
When evaluating any verification tool, ask specifically which of the three layers it runs and whether it reports them separately or blends them into one combined confidence score. A blended score can obscure important information: a citation that fails proposition-support but passes existence and quote-match might still receive a superficially reassuring overall grade, when in fact the most consequential risk (citing a real case for the wrong point) has gone undetected.
Related reading
- How to Check for AI Hallucinations in Legal Briefs
- The Hidden Risk of AI-Generated Quotations in Legal Writing
A common question
Can I run all three checks with one free tool?
CourtListener covers existence and quote-match well through its citation lookup and full-text search. Proposition-support generally requires actually reading the relevant passage yourself, since no free automated tool reliably replaces human legal judgment on that specific question today.
Keep a simple written note of which of the three checks you performed for each citation and when; it costs seconds and gives you a defensible record if your reasonable-inquiry process is ever questioned by a court or a client.
None of these three checks require a law degree to perform mechanically, which is exactly why delegating the mechanical portions to trained non-attorney staff, while reserving proposition-support judgment for a supervising attorney, can be an efficient division of labor for busier practices.
Revisit this framework every time a new AI tool is adopted in your practice, since different tools fail in different proportions across the three layers.