Yes — ChatGPT and other general-purpose large language models can and do cite fake cases, complete with plausible-sounding case names, realistic docket numbers, and quoted passages that read like real judicial writing. This is not a rare glitch; it is a predictable consequence of how these models generate text.
Why it happens
A general-purpose language model generates text by predicting statistically plausible continuations, not by retrieving verified facts from a legal database (unless it has been specifically connected to one with retrieval-augmented generation). When asked for a case supporting a given proposition, a model without live database access will often generate a citation that has the correct shape and style of a real case, because that is what it has learned real citations look like — without any underlying guarantee the specific case exists.
How often this happens
The Stanford RegLab/HAI study "Hallucinating Law" (Dahl, Magesh, Suzgun & Ho, 2024) found general-purpose models hallucinate on legal queries at rates between 58% and 82%, depending on the model and question type. This is dramatically higher than most users expect, especially given how confident and well-formatted the fabricated output typically looks.
The Mata v. Avianca example
In Mata v. Avianca, Inc., No. 22-cv-1461 (S.D.N.Y. 2023), ChatGPT produced multiple fake cases, including invented docket numbers and quoted excerpts, that were incorporated into a filed brief. When the attorney asked ChatGPT to confirm the cases were real, the model reaffirmed that they were — a detail from the record that underscores why asking the AI tool itself to double-check its own work is not a reliable verification method.
Does this happen with legal-specific tools too?
Yes, at a lower but still meaningful rate. A Stanford follow-up benchmark tested purpose-built legal RAG products (Lexis+ AI, Westlaw AI-Assisted Research, Ask Practical Law AI) and found hallucination rates in the roughly 17% to 34% range depending on the product. Purpose-built tools with database access reduce the problem; they do not eliminate it.
Expert perspective
Researchers on the Stanford study have specifically warned against treating model confidence as a proxy for accuracy: a fabricated case is typically presented with the same fluent, assertive tone as a real one, with no built-in signal distinguishing the two. That is precisely why independent verification against a primary source is necessary regardless of how confident the output sounds.
How to catch it every time
- Check every case citation for existence against a primary source like CourtListener, not against the AI tool's own reassurance.
- Check every quotation word-for-word against the actual opinion text.
- Confirm the case supports the specific proposition it is cited for.
- Never accept "yes, that case is real" from the same tool that generated the citation as sufficient verification.
A common question
If I ask ChatGPT to double-check its own citations, is that enough?
No. The Mata record shows the same model reaffirmed fabricated cases as real when asked directly. Verification needs an independent primary source, not a second query to the same tool.
Related reading
- AI-Hallucinated Case Law: Real Examples From the Public Record
- How to Verify Legal Citations From ChatGPT
- Which Legal AI Tools Hallucinate Least?
- Mata v. Avianca: The Lessons Two Years Later
- How to Check for AI Hallucinations in Legal Briefs
Check a brief before you file it →
Why ChatGPT specifically became the recurring example
ChatGPT's prominence in this conversation owes largely to its early, dominant market position among general-purpose AI tools rather than any evidence it hallucinates meaningfully more than comparable general-purpose models. The Stanford RegLab/HAI research measuring 58% to 82% hallucination rates on legal queries examined multiple general-purpose models, not ChatGPT alone, and found the underlying pattern — confident generation without database-verified grounding — is common across the category, not unique to any single vendor's product.
Does asking for sources help?
Asking a general-purpose model to "cite your sources" typically produces a citation formatted as if it were sourced, without necessarily reflecting an actual retrieval step behind the scenes, since the model may simply generate a plausible-looking source reference using the same generative process that produced the original claim. This is meaningfully different from a retrieval-augmented tool that genuinely queries a case-law database and returns a source it can point back to, which is a real, verifiable distinction worth understanding before trusting either kind of citation.
What responsible use actually looks like
Using ChatGPT or a similar tool to brainstorm arguments, outline a structure, or explain a legal concept in plain language carries relatively low risk, since none of that output is typically filed verbatim. The risk concentrates specifically at the point where AI-generated citations or quotations make it into a filed document without independent verification against a primary legal source.
Final takeaway
Yes, ChatGPT cites fake cases, and no, this is not a defect specific to one company's product that switching tools alone will solve. Independent verification against a primary source is the only reliable fix regardless of which general-purpose AI tool you use.
A common question
Are newer versions of ChatGPT less likely to hallucinate citations?
Model updates have generally improved factual accuracy on many tasks, but the Stanford RegLab/HAI research and its follow-ups have found hallucination on legal-specific queries remains meaningful across model generations, since general-purpose models are still not grounded in a verified legal database by default. Treat every new model version with the same verification discipline as the last one until independent benchmarking says otherwise.
Related reading
The safest working assumption for any AI tool, present or future, general-purpose or legal-specific, is that some nonzero hallucination rate exists until independently verified otherwise on your specific query.
Share this article with anyone in your office who uses AI tools casually for research, even informally, since the risk applies equally regardless of how the output is ultimately used.
A five-minute conversation now costs far less than a sanctions order later, regardless of how confident everyone in the room currently feels about their own AI usage habits.