Skip to main content
Citation Safe

From the makers of Citation Safe

The Hallucination Playbook for Lawyers

A practical guide to AI citation hallucination in legal filings — how it happens, what courts are doing about it, and how to build a verification workflow before you file, not after a judge or opposing counsel does it for you. Grounded in our own sanctions database and the same three-layer verification model that powers the product.

Free web edition, published as written

All ten chapters are complete below. This is the free web edition of the playbook, not a preorder for a finished book: there is nothing to buy here and nothing gated behind a purchase. If we add further chapters or revise these, the notify form below tells you when.

Get the PDF edition

Same ten chapters, formatted for printing or saving as a PDF from your browser — no manuscript to buy, nothing gated beyond an email address so we know who to notify about future revisions.

Chapter 1. How AI Hallucinates a Citation

Ask a large language model for a case citation and it will give you one. That is the entire problem in one sentence. A language model does not look anything up when it writes a citation — it predicts the next most statistically likely token, one at a time, based on patterns learned from the text it was trained on. Real citations follow a rigid, learnable format: a case name, a reporter volume, a page number, a court, a year. That format is exactly the kind of pattern a language model is very good at reproducing. What it is not good at — what it structurally cannot do on its own — is checking whether the specific case, volume, and page it just generated corresponds to something that actually exists in a real reporter.

This is why AI-hallucinated citations are dangerous in a way ordinary human error is not. A tired associate who misremembers a case name usually produces something that looks a little off — a wrong year, an unusual reporter abbreviation, a citation that does not quite parse. A language model does not make that kind of mistake. It produces a citation that is perfectly formatted, plausible in context, and stylistically indistinguishable from a real one, because it was trained on millions of real ones. The defect is invisible to the eye. You cannot proofread your way out of it, because proofreading checks whether something looks right, and a hallucinated citation looks exactly right. The only way to catch it is to check it against the source it claims to come from — a lookup, not a read-through.

Citation Safe’s own public sanctions database, built from Damien Charlotin’s court-record dataset of AI hallucination cases, is a good way to see this pattern in the wild rather than in the abstract. Take Akerlund v. Atlas Air, Inc., et al., an Eleventh Circuit matter decided July 10, 2026: the court’s sanction order lists case after case that counsel cited as authority, each one found by the court to be a nonexistent, hallucinated citation. Not one obviously wrong citation — a pattern of them, spread across a filing, each formatted like a normal case cite because that is what the model producing them was optimized to do. That is the first and most common failure mode in the taxonomy this playbook uses throughout: Fabricated. The cited case simply does not exist.

A second, distinct failure mode shows up in Aragon v. Industrial Claim Appeals Office, a Colorado matter decided July 2, 2026, with a Warning as its outcome. There, the underlying case was real — Pinkstaff v. Black & Decker exists — but counsel attributed a quotation to it that the court found does not appear in that decision, in a case the court further noted is not even a workers’ compensation matter. This is the second failure mode: False Quotes. The citation exists, but the language attributed to it does not, or does not say what the filing claims it says. This is arguably more dangerous than outright fabrication, because a reader who checks only that the case exists will conclude everything is fine. The case is real. The quote is not.

A third failure mode, Misrepresented, is different again: the case exists, the quoted language (if any) is accurate, but the holding does not actually support the point it is cited for. A real case, cited honestly on its face, pointed at the wrong proposition. This is the hardest of the three to catch by lookup alone, because nothing about the citation itself is false — only its use. A fourth pattern in the same database, Outdated Advice, covers a related but separate problem: authority that was once good law and is cited as if it still is, when it has since been reversed, superseded, or withdrawn.

Notice what all four failure modes have in common: none of them are things a spell-checker, a citation-formatter, or a careful second read of your own brief will reliably catch. Fabricated and False Quotes citations are specifically built to survive a human read-through, because the model that produced them was trained to produce exactly the kind of text a human read-through is designed to approve. Misrepresented and Outdated Advice failures require going back to the actual source and checking what it actually holds and whether it still holds it — work a reviewing attorney can do, but work a hallucinating drafting tool has already, by construction, failed to do for you.

This is also why “we have a policy against using AI” is not a durable answer, and why a growing number of federal and state courts have begun issuing standing orders that specifically address generative AI use in filings — some requiring a certification, some simply putting counsel on notice that hallucinated citations will be treated as a sanctionable failure to conduct reasonable inquiry. Whether or not your firm has an official AI policy, the citations in front of you may have been touched by a tool that produces exactly this class of error, and the only response that actually addresses the risk is checking the citations themselves against their claimed sources — not trusting that they were checked because they look right.

That is the shape of the rest of this playbook. Chapter 2 breaks down the four failure modes above in more detail, using the exact taxonomy applied across every teardown in Citation Safe’s sanctions database. Chapter 5 covers the three-layer verification model — existence, quote-match, and proposition support — built specifically to catch each of these failure modes at the layer where it actually occurs, rather than asking another language model whether a citation “looks” correct, which does not solve the underlying problem this chapter just described. Read the two real teardowns cited above in full at /blog/sanctions before you next rely on AI-assisted drafting in a filing — the actual court language is more persuasive than any summary of it, including this one.

Chapter 2. The Three Failure Modes That Are Actually Four

Chapter 1 named four failure modes without spending much time on any one of them: Fabricated, False Quotes, Misrepresented, and Outdated Advice. These labels are not something Citation Safe invented for marketing. They come directly from the taxonomy embedded in Damien Charlotin’s independently maintained, CC0-licensed database of AI hallucination cases, the same public dataset that underlies every teardown published in Citation Safe’s own sanctions database. Treating “the citation was wrong” as one problem is how a verification process ends up checking for the wrong thing. Treating it as four distinct problems, each with its own signature and its own detection method, is how a process actually catches them.

Fabricated is the simplest to define and, in the public record, the most common: the cited case does not exist at all, under that name, in that reporter, at that volume and page. Akerlund v. Atlas Air, Inc., et al., the Eleventh Circuit matter from Chapter 1, is a fabrication case in its purest form: a list of citations, each formatted like a normal case cite, none of them corresponding to an actual decision. A fabrication is caught by a single kind of check: does an authoritative index (for case law, a database of published opinions; for a statute, the text of the code itself) have a record matching what was cited. No reading comprehension is required. It is a lookup.

False Quotes is a different animal, because the case being cited is real. Aragon v. Industrial Claim Appeals Office, the Colorado matter also from Chapter 1, shows the pattern: Pinkstaff v. Black & Decker exists, but the quotation attributed to it does not appear in that decision, and the court further noted the case is not even a workers’ compensation matter, meaning the quote was not just inaccurate but attached to a case with no plausible connection to the point it was cited for. A false quote survives a check that only confirms the case exists, because the case does exist. Catching it requires a second, independent check: pulling the actual text of the source and searching it for the quoted language, not for a paraphrase or a close match, for the words themselves.

Misrepresented is harder still, because both the case and the quotation (if any) can be entirely accurate, and the citation can still be wrong. A real case, quoted correctly, cited for a legal proposition it does not actually stand for, passes both an existence check and a quote-match check. Nothing about the citation itself is false. Only its use is. This is the failure mode a citation-checking tool built purely on lookups cannot catch, because a lookup has nothing to compare the proposition against. Catching a misrepresentation requires reading the source and asking whether its holding actually supports the sentence it was attached to, a genuinely different kind of check than the first two.

Outdated Advice is the fourth pattern, and it is easy to mistake for one of the other three because on its face the citation is completely accurate: the case existed, the quote is real, and at one point it stood for exactly what the brief says it stands for. The defect is time. A holding that has since been reversed on appeal, superseded by statute, or overruled by a later decision is still, mechanically, a real case with a real quote. Nothing about the citation format signals its own obsolescence. Catching this requires checking not just whether authority exists, but whether it is still good law as of the filing date, a currency question rather than an existence question.

Notice that these four failure modes sort naturally into two different kinds of checks. Fabricated and False Quotes are deterministic: a lookup either finds the case or it does not, and the quoted text either appears in the source or it does not. There is no judgment call involved, which means these two checks can be run with zero risk of a false positive introduced by an LLM guessing. Misrepresented and Outdated Advice are not purely mechanical: they require some form of comparison between what the source actually says and what the filing claims it says, or between when a holding was decided and whether anything since has changed its force. Chapter 5 covers exactly how Citation Safe’s own verification engine is built around this split: existence and quote-match as fully deterministic layers, proposition support as an assisted, advisory layer that is never phrased as a final answer.

The practical takeaway for this chapter is narrower than “check your citations.” It is: know which of these four problems you are actually checking for before you decide a citation-checking process is working. A process that only confirms a case exists will pass every False Quotes and Misrepresented citation without comment, and will do so with total confidence, because from that process’s point of view, nothing is wrong. The next chapter walks through how the public sanctions database records real instances of exactly these four patterns, and how to search it for your own jurisdiction before you assume a pattern that shows up elsewhere would not show up in a filing in front of a judge you appear before.

Chapter takeaways

  • The exact category labels used in every Citation Safe teardown — Fabricated, False Quotes, Misrepresented, and Outdated Advice — and what each one means procedurally.
  • Why Fabricated and False Quotes are the two patterns that show up most often in the public sanctions database.
  • Why "the case is real" and "the quote is real" are two different questions that require two different checks.

What to do today

  • Pull three recent AI-assisted filings from your own practice and check whether you can independently confirm the case name and the quoted language for every citation in each.
  • Filter /sanctions-database to your jurisdiction to see which failure modes have actually shown up locally.

Chapter 3. Inside the Sanctions Database

The public sanctions database at /sanctions-database exists because reading about hallucinated citations in the abstract is much less persuasive than reading an actual order. Every entry traces back to Damien Charlotin’s independently maintained, CC0-licensed AI hallucination-cases dataset. Citation Safe does not compile these cases itself; it ingests the public dataset and builds a structured teardown around each entry using the case’s own recorded fields, with no model call and no editorial interpretation involved in generating the teardown text.

The procedural path in most of these cases looks similar across jurisdictions, even though the underlying facts vary. It starts when opposing counsel, or occasionally the court’s own chambers staff, cannot locate a cited case. In Mata v. Avianca, Inc., 678 F. Supp. 3d 443 (S.D.N.Y. 2023), the case that effectively started this entire area of judicial attention, Avianca’s lawyers told the court they had been unable to locate several cases cited in the plaintiff’s brief. The court could not locate them either. Judge P. Kevin Castel ordered plaintiff’s counsel to produce copies of the cited cases, and what came back did not hold up: several of the “cases” did not exist, and at least one existed in the record only because the AI tool that generated it had, when asked to confirm its own citation, replied that it could be found in reputable databases. On June 22, 2023, the court dismissed the underlying personal-injury claim and imposed a $5,000 sanction, finding that counsel had acted with subjective bad faith sufficient to support sanctions under Federal Rule of Civil Procedure 11.

That sequence, flag a citation, ask counsel to produce it, evaluate what comes back, is the pattern that repeats across the sanctions database, whether the outcome is a formal Rule 11 sanction, a warning, or, less often, no adverse finding once counsel corrects the record before the court rules. The database’s outcome field records which of these actually happened in each case, and it is worth reading a few entries in full rather than assuming every entry ends the same way. Aragon v. Industrial Claim Appeals Office, for instance, resulted in a warning rather than a monetary sanction, a materially different outcome than Akerlund’s.

Searching the database works across three axes: jurisdiction (state, or federal circuit), the AI tool identified in the record (where the record identifies one; Charlotin’s underlying schema deliberately records “unidentified” or “implied” rather than guessing a tool when a court order or reporting does not name one explicitly), and year. This matters practically because the exposure is not evenly distributed. A search filtered to your own jurisdiction and the last two years is a five-minute read that tells you whether this has already happened somewhere you, or your opposing counsel, might file next.

The AI-tool field is worth a second look on its own, separate from jurisdiction and year, because it is the field most likely to be incomplete rather than wrong. Charlotin’s schema records a named tool only when a court order, a brief, or reporting on the matter states it explicitly; matters where AI use is evident from the pattern of fabricated citations but no named product appears in the record are tagged “unidentified” or “implied” rather than assigned a guess. Do not read a heavy concentration of entries naming one particular tool as evidence that tool is uniquely prone to hallucination; it may just as easily reflect which tool happened to be named in the underlying record for cases decided so far. The safer inference from the database is about the pattern of failure, not a ranking of which AI product is safest to trust.

It is worth being precise about what this database is not. It is not, and does not claim to be, a complete national record of every AI-hallucination sanction ever issued. It tracks publicly reported cases: cases that produced a written order, a published opinion, or reporting substantial enough to be captured in the underlying dataset. Sanctions handled informally, off the record, or in jurisdictions with less robust public court-record access will not appear, even if they happened. Treat an empty search result for your jurisdiction as “nothing publicly reported yet,” not as “this cannot happen here.”

The database also is not, and does not present itself as, legal advice or a certification that any specific brief is safe. It is a record of what has already gone wrong for other filers, organized so the pattern is visible rather than buried across individual court dockets. The purpose of reading it before you rely on AI-assisted drafting is the same purpose a compliance officer reads enforcement actions before signing off on a new process: not because the exact fact pattern will repeat, but because the failure mode will.

The next chapter turns to what courts have started doing about this pattern at the level of standing orders and local rules, starting with the order that effectively started this entire category of judicial response.

Chapter takeaways

  • The procedural path from a flagged citation to a show-cause order to a published sanction or warning.
  • How to search the database by jurisdiction, AI tool, and year at /sanctions-database.
  • That the database tracks publicly reported cases — it is informational, not a complete national record.

What to do today

  • Search /sanctions-database for your own jurisdiction and read the most recent case in full.
  • Do not assume "it hasn't happened here" without checking first.

Chapter 4. Court Standing Orders and AI Disclosure

Mata v. Avianca did more than produce a widely reported sanction. It triggered, within days, the first judge-specific standing order anywhere in the federal system addressing generative AI in court filings. On May 30, 2023, Judge Brantley Starr of the U.S. District Court for the Northern District of Texas issued a “Mandatory Certification Regarding Generative Artificial Intelligence,” requiring that all attorneys and pro se litigants appearing before him file, together with their notice of appearance, a certificate attesting either that no portion of any filing was drafted using generative AI, or that any language an AI tool did draft was checked for accuracy by a human being using print reporters or traditional legal databases. Judge Starr’s order named ChatGPT, Harvey.AI, and Google’s Bard as examples of the tools it covered, and it was explicit that failing to file the certificate could result in the filing being struck.

The certification requirement also reaches a question standing orders in this area do not always spell out explicitly: supervision. A partner who signs a brief a junior associate drafted with AI assistance is not insulated from a certification’s requirements by having delegated the drafting. ABA Formal Opinion 512, discussed below, treats supervisory responsibility for subordinate lawyers’ and nonlawyer assistants’ use of generative AI as its own distinct duty, separate from the supervising lawyer’s own direct use of the tools. A firm’s policy on AI-assisted drafting is incomplete if it only addresses what the person typing the prompt has to do and says nothing about what the person reviewing and signing the final product is independently responsible for confirming.

Judge Starr’s order was the first, not the only. Within months, other judges in other districts issued their own standing orders, some closely modeled on his certification language, others taking a lighter touch: simply putting counsel on notice that the court is aware generative AI tools produce plausible-looking fabrications and that ordinary Rule 11 obligations apply in full. The result today is a patchwork rather than a single uniform national rule: some courts require an affirmative, on-the-record certification before a lawyer’s first filing; others have no AI-specific order at all and rely entirely on existing rules of professional conduct and Rule 11 to do the same work.

This unevenness matters practically. A lawyer who checks whether their district or their specific judge has published a standing order, and finds none, should not conclude that AI-drafted citations carry less risk there. It means only that the court has not chosen to create a separate, AI-specific paper trail for an obligation that exists regardless: a signature on a filing is already, under Rule 11, a certification that a reasonable inquiry was made into the factual and legal contentions in it. A hallucinated citation does not become more or less sanctionable because a specific standing order does or does not exist; it becomes easier to prove counsel violated an obligation they were on express, individualized notice of.

The other major development in this area is not judicial but professional: on July 29, 2024, the American Bar Association’s Standing Committee on Ethics and Professional Responsibility issued Formal Opinion 512, “Generative Artificial Intelligence Tools,” its first formal ethics guidance addressing how the Model Rules of Professional Conduct apply to a lawyer’s use of generative AI. The roughly fifteen-page opinion does not create new rules; it maps existing duties, including competence, confidentiality, communication with clients, candor to the tribunal, and supervision of subordinate lawyers and nonlawyer assistants, onto the specific ways generative AI tools are actually used in practice. Its position on competence is consistent with what every reported sanctions case in this area already demonstrates: the degree of independent verification required depends on the tool and the task, and reviewing AI-drafted citations for accuracy is not optional simply because a tool produced them quickly.

Put together, the standing-order patchwork and the ABA’s ethics guidance point in the same direction: the paper-trail question (“did I sign a certification”) is secondary to the substantive question (“did I actually verify what I filed”). A signed certification that a human checked every AI-drafted citation is not itself a verification; it is a representation that one happened. If it did not happen, or happened by skimming rather than by checking each citation against its claimed source, the certification does not protect against a Rule 11 problem. It becomes one, because a false certification is its own separate misrepresentation to the court.

The practical response is the same whether or not your court has published a standing order: build a verification step you can actually describe in one sentence, one that runs before a document is filed, not as a reaction to an order to show cause. What that step looks like in practice, sequenced across every citation in a document rather than just the ones that feel uncertain, is the subject of Chapter 7. First, Chapter 5 lays out the three-layer model this playbook keeps referring back to, because a verification step is only as good as knowing exactly what each layer of a check does and does not confirm.

Chapter takeaways

  • Why a growing number of federal and state courts have adopted standing orders addressing generative AI use in filings.
  • That these orders vary — some require an affirmative certification, others simply put counsel on notice.
  • That a hallucinated citation can be treated as a failure to conduct reasonable inquiry whether or not a specific AI standing order applies to your filing.

What to do today

  • Check whether the court(s) you file in have published an AI-specific standing order or local rule.
  • If one exists, confirm your firm's actual filing process satisfies it today — not that someone read it once.

Chapter 5. The Three-Layer Verification Model

Every failure mode from Chapter 2 maps onto a specific layer of the same three-layer check, which is the actual verification engine behind Citation Safe (and, with a different source adapter per vertical, Tax Cite Safe, Med Cite Safe, and FAR Check). The three layers are Existence, Quote Match, and Proposition Support, described in full at /methodology, and it matters which of these three produced a given check, because they are not equally reliable and they are not built the same way.

Layer 1, Existence, is a deterministic lookup: does this citation exist in the authoritative public source for its type. For case law, that means checking against CourtListener’s database of published opinions (the Free Law Project’s public repository), including, where the citation type supports it, a name cross-check confirming the case name in the filing actually matches the case sitting at that reporter citation. No LLM is involved in producing this verdict. A real reporter slot occupied by a different case does not pass; the check is not “does something exist here” but “does the specific thing claimed exist here.” Layer 1 is what catches Fabricated citations, the failure mode described in Chapter 2, because a fabricated case simply will not be found.

Layer 2, Quote Match, is also fully deterministic, and it answers a different question: if the filing quotes the source, does that exact language actually appear in the source’s text. This is a text search, not a paraphrase-similarity comparison, which is deliberate: a paraphrase that captures the gist of a real passage is a different problem than a fabricated quotation, and conflating them would blur exactly the distinction Chapter 2 draws between a real case correctly quoted and a real case falsely quoted. Layer 2 is what catches False Quotes, and it catches it independently of Layer 1, meaning a citation can pass Layer 1 (the case is real) and still fail Layer 2 (the quote is not).

Layer 3, Proposition Support, is fundamentally different from the first two, and the difference is disclosed rather than hidden: it is LLM-assisted, it is advisory, and it is reported with a confidence figure rather than a flat pass or fail. It asks whether the cited authority actually appears to support the proposition it is attached to, the question behind the Misrepresented failure mode from Chapter 2. This is not a question a pure lookup can answer, because the case can be entirely real and the quote entirely accurate while the citation is still wrong. Because Layer 3 depends on a live LLM connection and involves judgment rather than a lookup, it is never phrased as “verified correct.” It is evidence for a lawyer’s own read of the source, not a substitute for reading the source. When the connection Layer 3 depends on is unavailable, the honest response is to not run it and say so, rather than show a stale or guessed result; live status on whether Layer 3 is currently running is published, not asserted once and left stale, at /quality.

A confidence figure attached to a Layer 3 result deserves a specific caution: it is not a probability that the citation is correct in any statistically calibrated sense, and it should not be read as one. It is a reported measure of how strongly the model’s pass over the source text supports the proposition, produced by the same kind of system that produces the hallucinations this entire playbook is about. Layer 3 is useful precisely because it is disclosed as an assisted, imperfect signal rather than dressed up as a verdict, and the same discipline that applies to reading any AI output applies here: treat a high number as a reason to look closer with more confidence you will find support, not as a reason to skip looking.

Outdated Advice, the fourth failure mode from Chapter 2, does not map cleanly onto any single layer above, because it is fundamentally a currency question rather than an existence, quote, or proposition question. A case can pass all three layers today and still have been reversed, superseded, or withdrawn since it was decided, and none of Existence, Quote Match, or Proposition Support as described above is built to detect that on its own. This is a disclosed limitation rather than a gap papered over with a vague claim of coverage, and knowing that boundary is itself part of using any verification tool responsibly.

The reason to care which layer produced a given verdict, rather than just whether a check “passed,” is that the layers carry genuinely different reliability. A Layer 1 or Layer 2 failure is a fact: the source either has the record or it does not, the text either contains the quoted words or it does not. A Layer 3 result is a probabilistic judgment about a genuinely hard question, dressed in a confidence figure rather than a certainty. Chapter 6 goes through exactly what each possible verdict, VERIFIED, UNCONFIRMED, NOT FOUND, RETRACTED, and SOURCE UNAVAILABLE, actually asserts, because the practical risk in adopting any citation-checking tool, this one included, is treating every passing result the same way regardless of which layer produced it.

Chapter takeaways

  • Layer 1, Existence — a deterministic lookup against the authoritative public source. No LLM is involved in this verdict.
  • Layer 2, Quote match — a deterministic check of whether quoted language actually appears in the source's text.
  • Layer 3, Proposition support — LLM-assisted and advisory, reported with a confidence figure, never phrased as "verified correct."

What to do today

  • Read /methodology in full so you know exactly which layer would have caught which failure mode from Chapter 2.
  • Check /quality for the live, measured error rate per layer before deciding how much weight to give Layer 3 output today.

Chapter 6. Reading a Verdict

A citation check is only useful if the person reading its output knows exactly what each possible result is claiming, and exactly what it is not claiming. This chapter walks through the specific verdict language behind every check, drawn directly from the single language-contract module that generates this copy everywhere it appears, so the same word means the same thing on every report, every time.

VERIFIED means the citation exists in the named source as of the named timestamp. Where the citation type supports it, this includes the name cross-check described in Chapter 5: the case name the filing claims has to match the case actually occupying that reporter citation. A real reporter slot occupied by a different case does not produce VERIFIED, because “a citation exists at this address” and “the citation you wrote is the one at this address” are different claims, and only the second one is safe to rely on.

NOT FOUND and UNCONFIRMED sound similar and mean substantially different things, and the difference matters more than almost any other distinction in this chapter. NOT FOUND means the authoritative index for that citation type was reachable and affirmatively has no record of it: a confident finding of non-existence. UNCONFIRMED means the citation could not be confirmed as cited, but the signal behind that failure is not confident non-existence; it might be a claimed-name mismatch, or a case where only a weak fuzzy search was possible rather than an exact lookup. Treating UNCONFIRMED as a soft pass, “probably fine, just didn’t quite match,” is exactly the mistake this distinction exists to prevent. UNCONFIRMED is an instruction to go look the citation up yourself, not a lesser form of good news.

RETRACTED is never folded into plain VERIFIED, even though a retracted source technically still exists at the citation given. A retraction is itself an authoritative signal, whether that is a court’s own subsequent order or, in the medical-citation context covered in Chapter 9, a journal’s or Crossref’s retraction flag, and collapsing it into a plain VERIFIED would hide the single most important fact a reader needs about that source.

SOURCE UNAVAILABLE is an honesty mechanism, not a verdict about the citation itself. It means the authoritative source could not be reached at check time, full stop, and the result reports that outage rather than guessing in either direction. A tool under pressure to always return a clean answer has an incentive to quietly treat an unreachable source as a pass; the discipline here is to say “we could not check this” out loud instead, even though that is a less satisfying thing to tell a paying customer than a green checkmark.

What to actually do with a SOURCE UNAVAILABLE result is worth spelling out, because leaving it unresolved is not a safe default. Treat it the same way you would treat a citation you have not checked at all: confirm it manually against the source directly (a docket search, a database your firm already subscribes to, or the source’s own site) before the document is filed, and retry the automated check separately, since an outage at check time does not mean the source will stay unreachable. A filing that goes out with an unresolved SOURCE UNAVAILABLE result on a load-bearing citation has not been verified; it has been told, honestly, that verification did not happen yet.

Quote-match verdicts follow the identical discipline in miniature: QUOTE CONFIRMED, QUOTE UNCONFIRMED (worded as “we could not find this language,” never as “you misquoted,” because a failure to locate a passage is not proof the passage does not exist elsewhere in a large document), SOURCE UNAVAILABLE, or NO QUOTE ATTACHED when there was nothing to check in the first place.

One more distinction is worth holding onto across all of these verdicts: a check reporting VERIFIED, NOT FOUND, or RETRACTED is telling you something about the source. A check reporting UNCONFIRMED or SOURCE UNAVAILABLE is telling you something about the check itself, that it could not reach a confident answer this time, for reasons that may have nothing to do with whether the citation is good. Conflating the two categories, source facts and check limitations, is the single easiest way to misread an otherwise accurate report.

The single habit this chapter is trying to build is simple to state and easy to skip under deadline pressure: before relying on the output of any automated citation check, this one or anyone else’s, know exactly what its passing result is asserting. “Passed” from a tool that only runs an existence check is a narrower claim than “passed” from a tool that also checked the quote and the proposition, and a filing built on the assumption that a green result means “this citation is completely safe to rely on” is exactly the assumption that turns a tool limitation into a sanctions exposure. Chapter 7 turns this into a concrete, repeatable sequence: where verification actually belongs in a filing’s drafting process, and what happens when a result looks wrong.

Chapter takeaways

  • What VERIFIED actually asserts, including the name cross-check applied to case citations.
  • The difference between NOT FOUND (confident non-existence) and UNCONFIRMED (inconclusive, not a negative finding).
  • Why RETRACTED is never folded into plain VERIFIED, and what SOURCE UNAVAILABLE means versus a real result.

What to do today

  • Before relying on any automated citation check — ours or anyone else's — confirm you know exactly what its "pass" result is asserting.
  • Treat UNCONFIRMED as "go look this up yourself," not as a soft pass.

Chapter 7. Building a Pre-Filing Verification Workflow

Every chapter so far has described a piece of the problem or a piece of the check. This chapter is about sequencing: where, concretely, verification belongs in the life of a document, because the difference between a firm that has access to a citation checker and a firm that has a verification workflow is entirely about when the check happens and who is accountable for it happening.

The wrong place for verification, and the place it ends up by default absent a deliberate decision otherwise, is after opposing counsel’s motion, or after a court’s order to show cause. By that point the filing is already on the docket, the sanctions exposure already exists, and the check that should have happened before signature is now happening in a much more adversarial, much more public context. Every case in the sanctions database discussed in Chapter 3 is, procedurally, an example of verification happening too late: after a flag, not before a filing.

The right place is a fixed point in the drafting process that happens on every filing that cites authority, not just the ones that feel uncertain. This is a harder discipline than it sounds, because the citations that feel safest, the ones pulled from a memo the associate is confident about, from a brief that has been through several rounds of internal review, or from a source that “everyone knows” is right, are exactly the ones most likely to skip a check that is applied selectively rather than universally. An AI-hallucinated citation does not look uncertain. It looks exactly as confident as a real one, which is the entire point made in Chapter 1: proofreading and confidence-based triage cannot substitute for checking every citation against its claimed source.

In practice this means sequencing the check across the whole document, not spot-checking. Pull every case citation, every quoted passage attributed to a source, every statute or regulation number, and run each one through existence, quote-match, and (where available) proposition-support checks before the document leaves the building, meaning before it is filed or sent to opposing counsel or the court in any form, not after a final proofread that assumes the substance is already settled.

A dispute-and-recheck step is part of a real workflow, not an afterthought. When a verification result looks wrong, whether because it flags a citation the drafting attorney is personally confident about, or because a source appears temporarily unreachable, the correct response is to run the recheck through the same source the original check used, not to override the result based on personal confidence. Citation Safe’s own dispute mechanic works this way for exactly this reason: filing a dispute immediately re-runs the check against the same authoritative source and the same name cross-check where one applies, rather than escalating straight to a human override. Two asymmetric rules apply to what happens next, described in full at /refund: a re-check that confirms the original result was wrong (a paid check stamped VERIFIED that turns out not to exist) triggers an automatic refund with no negotiation, logged to the tool’s public error rate; a re-check that would upgrade a result to VERIFIED is never automatic, and is always escalated for a human to confirm, because a wrong downgrade only costs a lawyer a moment’s inconvenience double-checking a real citation, while a wrong upgrade risks a lawyer relying on a citation that is not actually safe. The asymmetry is deliberate, not an oversight.

Emergency and same-day filings are the hardest test of a workflow, precisely because they are the circumstance most likely to produce an argument for skipping the check “just this once.” A workflow built only for the ordinary case, with an unstated exception for genuine time pressure, will fail exactly when the pressure to file something that looks right, fast, is highest, which is also when a drafting tool is most likely to have been leaned on heavily. The fix is not to make the check slower to accommodate emergencies; it is to decide, in advance and in writing, that an emergency filing still gets the existence and quote-match layers run before signature, even if a full proposition-support review has to happen on an expedited basis afterward. Deciding this in the moment, under the exact pressure that makes shortcuts tempting, is how the exception quietly becomes the rule.

The single-sentence test for whether a firm actually has a workflow, rather than a tool it owns a license to, is this: can you say, in one sentence, the exact point in your drafting process where every citation in a document gets checked against its claimed source. If the honest answer is “whenever someone remembers” or “when it feels necessary,” it is not a workflow yet, regardless of what software sits on the desktop. Chapter 8 turns to the harder, related question this raises: given all of this, where does a lawyer’s own judgment still have to do work no deterministic check, and no LLM-assisted layer, can do for them.

Chapter takeaways

  • Why verification belongs before a brief is filed, not as a reaction to opposing counsel's motion.
  • How to sequence a check across every citation in a document, not only the ones that feel uncertain.
  • What a dispute-and-recheck step looks like when a result seems wrong.

What to do today

  • Pick one filing currently in your pipeline and run every citation in it through a check before it goes out.
  • Write down, in one sentence, the exact point in your drafting process where verification happens. If you can't answer that, it isn't a workflow yet.

Chapter 8. Where Human Judgment Still Matters

Nothing in this playbook argues that a verification tool, this one or any other, replaces a lawyer’s judgment. The opposite point, made directly in Citation Safe’s own methodology documentation, is that Layer 3’s proposition-support review is described as evidence for a lawyer’s judgment, not a substitute for it, and that framing is not a disclaimer bolted on for liability reasons. It reflects a real, structural limit on what any of these checks, deterministic or LLM-assisted, can actually evaluate.

Start with what the deterministic layers, Existence and Quote Match, genuinely cannot do. They can confirm a case exists and that quoted language appears in it. They cannot tell you whether that case’s reasoning actually fits the argument it has been slotted into, whether a factually distinguishable case has been cited past the point its holding reasonably extends, or whether a string of individually accurate citations, read together, misrepresents how settled or unsettled an area of law actually is. These are not edge cases; they are the ordinary, everyday work of legal analysis, and no lookup performs it.

Put concretely: imagine a brief that cites a real, correctly quoted appellate decision for the proposition that a particular contract clause is enforceable, where the cited case actually involved a materially different clause in a different industry, decided on reasoning tied closely to facts that do not carry over. Layer 1 confirms the case exists. Layer 2 confirms the quoted language is really there. Neither layer has any way of knowing that the fact pattern does not transfer, because neither layer is built to compare fact patterns at all; that is squarely a legal-judgment question, and it is the kind of gap a careful reviewing attorney closes by actually reading the cited opinion, not by trusting two green checkmarks that were never designed to answer it.

Layer 3’s proposition-support check reaches further into this territory than the first two layers, but it is built, deliberately, to stop short of replacing the analysis rather than to complete it. A confidence figure attached to a Layer 3 result is not a verdict; it is a prompt. A high-confidence Layer 3 result on a proposition the drafting attorney has not personally traced back to the source’s actual holding is not evidence the citation is safe. It is evidence that an automated pass over the text did not find an obvious mismatch, a narrower and more modest claim.

The asymmetric-recheck rule from Chapter 7, that a downgrade following confirmed fault happens automatically but an upgrade never does, exists for exactly this reason. Getting a wrong result more cautious costs a few minutes of a lawyer’s time confirming something that was actually fine. Getting a wrong result more confident, automatically, risks a lawyer relying on a citation an automated recheck merely failed to flag rather than affirmatively confirmed. The system is built to fail toward caution, not toward convenience, because the failure mode this entire playbook is about, an AI-hallucinated citation that looks completely normal, is itself a failure toward false confidence, and a verification tool that also fails toward false confidence under pressure has not actually solved the underlying problem. It has just moved it one layer down.

The practical discipline this creates is to treat every deterministic pass and every Layer 3 confidence figure as an input to your own read of the source, never as a replacement for having read it. This is not a hedge to avoid liability; it changes what the actual verification step looks like day to day. Existence and quote-match failures should be resolved by looking at the source directly, not by re-running the check and hoping for a different answer. A high Layer 3 confidence figure on a citation central to the argument’s success should prompt reading the actual holding, not skipping that read because a tool already said it looked fine.

Some filings lean more heavily than others on exactly the kind of judgment call no tool can make: whether a real, correctly quoted case’s reasoning genuinely extends to a materially different fact pattern, whether an area of law that looks settled from a string of accurate citations is actually more contested than the string suggests, whether an argument’s overall persuasiveness depends on a citation doing more work than its actual holding supports. Identifying which of your current filings depend most on that kind of judgment, and reviewing those personally rather than delegating review entirely to a passing verification report, is the single highest-value habit this chapter can leave you with.

Chapter 9 moves outside case law entirely, because the same hallucination risk, and the same layered response to it, shows up in tax authority, medical literature, and government-contracts citations, each with its own coverage boundaries worth knowing before you assume a citation type is covered.

Chapter takeaways

  • Why proposition-support review is described as evidence for your judgment, not a replacement for it.
  • Why a re-check that would upgrade a result to VERIFIED is never automatic — only a downgrade following confirmed fault is.
  • What deterministic checks are structurally unable to evaluate, such as whether a real case's reasoning genuinely fits your argument.

What to do today

  • Read every Layer 3 confidence figure as a prompt for your own read of the source, not a final answer.
  • Identify which of your current filings rely most heavily on judgment calls no tool can make, and review those first.

Chapter 9. Beyond Case Law: Tax, Medical, and Government Citations

Everything in this playbook so far has used case law as its running example, but the underlying problem, an AI tool generating a citation that is perfectly formatted and completely wrong, is not specific to case law. It shows up anywhere citations follow a learnable, rigid format: Internal Revenue Code sections, medical literature identifiers, and government-contracts clauses all qualify, and each one is checked against a different authoritative source with its own coverage boundary, detailed in full at /methodology.

Tax authority splits into three tiers of confidence, not one. IRC sections (26 U.S.C.) are checked against Cornell’s Legal Information Institute, and Treasury Regulations (26 CFR) against the eCFR, both strong, direct lookups against a canonical text. Tax Court opinions are checked against CourtListener, scoped specifically to the Tax Court, but this is a structurally weaker signal than ordinary case-law checking: the citation field is frequently empty for Tax Court memo opinions, which turns the check into a caption-and-snippet search rather than a precise citation lookup, and that lower confidence is disclosed on the result rather than hidden behind a flat pass. Revenue Rulings, Revenue Procedures, IRS Notices, and Private Letter Rulings sit outside coverage entirely today, labeled as such rather than given a fake pass.

Medical citations are checked against three separate live sources depending on identifier type: PubMed IDs against NCBI’s E-utilities, DOIs against Crossref, and clinical trial registrations against ClinicalTrials.gov. The check that matters most in this vertical, and the one generic citation tools miss entirely, is retraction status: the signal that matches whichever identifier was cited, PubMed’s own “Retracted Publication” flag for a PMID, or Crossref’s retraction metadata (with a “RETRACTED:” title-prefix fallback for cases where structured metadata is incomplete) for a DOI, so that a paper the source itself flags as retracted returns a RETRACTED verdict rather than being folded into a plain VERIFIED. In litigation terms, an expert report that cites a paper without noting it was later retracted is exactly the kind of gap opposing counsel checks for before a Daubert challenge, and it is worth checking it yourself first, before they do.

Government-contracts citations split similarly into a strong-signal category and a weaker one. FAR and DFARS clauses are checked directly against acquisition.gov’s published text, including a currency signal drawn from Federal Acquisition Circular data where available, meaning a cited clause revision that no longer matches the version currently in effect can be flagged rather than silently treated as current. GAO bid-protest decisions are a harder case: the citation format is recognized and the canonical gao.gov link is constructed, but GAO’s site blocks some automated retrieval, and when the source cannot actually be fetched, the honest result is SOURCE UNAVAILABLE rather than a fabricated verdict pointing either direction. A bid protest built on a GAO decision deserves the same discipline as a case citation: confirm it against the actual decision, not against a tool’s confident-sounding restatement of what the decision probably says.

The throughline across all three of these verticals is the same discipline described in Chapter 6 for case law: know exactly what a passing result is actually claiming, and know where coverage stops rather than assuming a tool checks everything it touches. A tax practitioner relying on a Revenue Ruling citation, a med-mal defense counsel relying on a bare journal citation without a PMID or DOI, or a government-contracts specialist relying on a GAO decision the underlying source blocks from automated retrieval are all, in their own vertical, in the same position as a litigator relying on an UNCONFIRMED case citation: the honest label is “not confirmed,” not “probably fine.”

The coverage differences across these three verticals are not arbitrary; they track how each field’s authoritative sources actually publish data. Cornell LII and the eCFR publish IRC and Treasury Regulations text in a stable, machine-checkable format, which is why tax-statute checks are as strong as case-law checks. PubMed, Crossref, and ClinicalTrials.gov were built from the start as structured research infrastructure with identifiers designed for exactly this kind of lookup, which is why medical citation checking, retraction status included, is comparably strong. GAO, by contrast, was not built as a machine-readable public API, and blocking some automated retrieval is a byproduct of that design rather than a policy decision aimed at verification tools specifically. None of this is a reason to trust an uncovered citation type more; it is a reason to expect coverage to keep expanding unevenly across verticals as sources change, rather than all at once.

If your own practice touches any of these three areas, the practical step is to check the live coverage map at /quality before assuming a citation type you rely on regularly is actually covered by whatever tool you use, this one or another. An honestly disclosed coverage gap is not a defect; treating an undisclosed gap as full coverage is the actual risk, in exactly the same way an AI tool’s confident, well-formatted fabrication is more dangerous than an obviously sloppy human error. Chapter 10 closes this playbook with what all nine chapters so far add up to at the level of an actual firm policy, not just an individual habit.

Chapter takeaways

  • How IRC sections and Treasury Regulations are checked against Cornell LII and the eCFR, and why Tax Court citations carry lower confidence.
  • How PMIDs, DOIs, and clinical trial registrations are checked, including the retraction signal drawn from PubMed and Crossref.
  • How FAR/DFARS clauses are checked against acquisition.gov, and why GAO bid-protest decisions sometimes return SOURCE UNAVAILABLE instead of a verdict.

What to do today

  • If your practice touches tax, medical, or government-contracts citations, check the coverage map at /quality before assuming a citation type is covered.
  • Treat "outside coverage" labels as an honest boundary, not a false negative.

Chapter 10. A Firm-Wide Verification Policy

Everything in this playbook, the four failure modes, the sanctions database, the standing orders, the three-layer model, verdict semantics, workflow sequencing, the limits of automated checks, and the coverage boundaries across tax, medical, and government citations, is only as useful as whether it survives contact with an actual filing deadline on a day when the associate who usually remembers to run the check is out sick. A policy that lives in one person’s memory is not a policy; it is a habit with a single point of failure, and single points of failure are exactly what a firm-wide policy exists to remove.

A real policy answers three questions in writing, not just in conversation: which filings require a citation verification step, at what point in the drafting process that step happens, and who is accountable for confirming it happened before the document goes out under the firm’s name. “Before signature, not after” is the right answer to the second question for the overwhelming majority of filings that cite any outside authority, because every case in the sanctions database discussed in Chapter 3 involves a check that, if it happened at all, happened after a court or opposing counsel already flagged the problem rather than before the document was signed.

The dispute-and-refund mechanism described in Chapter 7 is worth building into a firm’s policy explicitly, not just relying on as a backstop, because it changes the cost calculus of running a paid verification check at all: if a paid check stamps a citation VERIFIED and that turns out to be wrong, the fee for that specific verification is refunded automatically once the fault is confirmed, no negotiation required beyond filing the dispute. That guarantee, and its boundaries (it covers the fee for the affected verification, not consequential losses, and it does not apply to UNCONFIRMED, NOT FOUND, RETRACTED, or SOURCE UNAVAILABLE results, because those verdicts are the tool correctly declining to make a claim it cannot support, not an error to refund), is spelled out in full at /refund, and a firm’s internal policy should point to that page rather than restate it from memory, so the two never drift apart.

The upgrade-never-automatic rule from Chapter 7 deserves its own line in any firm policy, separate from the refund mechanism: no automated recheck should ever be treated, by policy, as sufficient grounds to move a citation from a flagged or uncertain status to fully relied-upon without a named human confirming it. This is not a limitation to work around; it is the correct, conservative default, and a policy that quietly allows an automated upgrade to substitute for a human confirmation has reintroduced, at the firm-process level, exactly the false-confidence risk this entire playbook is about.

Assigning a named owner for the policy, not “the associate who currently knows how to do this,” is the detail that determines whether a policy survives staff turnover. A verification requirement that depends on institutional memory held by one person is fragile in exactly the way an unwritten filing-compliance rule of any other kind is fragile: it works until that person is unavailable, and it fails silently, because nobody else in the process knows the step was skipped until a court or opposing counsel points it out.

The accountable owner named above should not be a purely ceremonial role. Their job includes confirming, periodically rather than only when something goes wrong, that the verification step is actually happening as designed across the filings the policy covers, not merely that everyone agrees in principle it should be. A policy also needs a way to onboard new people into it and a way to notice when it stops working. New associates and lateral hires should be told the verification requirement as part of arriving at the firm, not left to absorb it informally from whoever happens to mentor them, and the policy should specify who delivers that briefing. Periodically checking the sanctions database discussed in Chapter 3 for new entries in your own jurisdiction is a reasonable, low-effort way to confirm the policy is still tracking a live risk rather than one the firm addressed once and stopped thinking about; a policy that is never revisited ages the same way an unread compliance manual does.

None of this requires a large process. It requires three things written down: which filings trigger the requirement, the specific point in drafting where it happens, and one named person accountable for it. That is the entire distance between a firm that has access to a citation-checking tool and a firm that has actually closed the gap this playbook opened with in Chapter 1: a language model that produces a perfectly formatted citation has no mechanism, on its own, for knowing whether that citation is real. The only thing standing between that structural fact and a sanctions order with your firm’s name on it is whether checking actually happens, every time, before a judge or opposing counsel does it for you instead.

Chapter takeaways

  • Why a policy that lives in one associate's memory is not a policy.
  • How a dispute-and-automatic-refund mechanism can work when a paid verification turns out wrong.
  • Why upgrades to VERIFIED always require human confirmation, even under an automated recheck.

What to do today

  • Write down, for your firm, exactly which filings require a citation check and at what stage — before signature, not after.
  • Assign a named owner for the policy so it survives staff turnover, the same way any other filing-compliance step does.

Get notified about new chapters

This page is the free web edition, published as chapters are written. We’ll email you when new chapters go up. No spam, no drip sequence.

Want the same discipline running on your filings today, not just in a book? See Citation Safe, how the verification engine works, real sanction teardowns sourced from the same database this playbook draws on, or the free verification tools.