Every firm that has read one of the sanctions horror stories quickly adds a line to the prompt: only cite real cases, don’t hallucinate, double-check every citation. This is called a citation directive, and it reads like an authoritative control. It isn’t one — and it fails hardest on the exact briefs where you were counting on it.
Many people are under the mistaken impression that an LLM is some sort of hypercomplex thinking machine with deterministic behavior. Over-simplifications aside, an LLM is just a giant probability map of words and patterns followed by other words and patterns. This stochastic system is not a database, it is not a search engine, and it is not a lookup table. It does not have the ability to know what is true or false. It can only produce the probabilistically likely next word in a sequence based on the patterns it has seen in its training data and assigned probabilities.
When you ask for citations, you're talking about deterministic results. LLMs are not determinististic and they have famously struggled with simple deterministic questions like "how many Rs are there in STRAWBERRY". LLMs aren't actually counting the Rs, just as LLMs aren't actually looking up cases in databases.
It's the same reason why an LLM can't even reliably count citations in a brief. Take a regular brief and count the citations manually. Let's say yours has 40 citations. Then ask your favorite frontier model LLM to identify all citations and output the count. Now run that prompt 1000x times. Sometimes it will identify all 40 citations, sometimes it will identify 38, sometimes it will identify 42. At the end of your 1000x runs, you'll have a whole distribution map of results because the LLM is not actually counting the citations. It's producing the next likely word based on assigned probabilities in patterns in its weighted token map. Even the best frontier models are still wrong on this simple task 20-40% of the time.
The LLM can't "only cite real cases" because there is no internal flag that says fabricated or authoritative.
Over-simplifications aside, every major LLM model has something called "completion pressure" -- a mathematically defined preference to produce an answer rather than tell you it can’t.
Adding a citation directive like “only cite real cases” doesn’t lift that pressure. Instead, it hands the model two orders that collide the moment real support runs out: reach the holding, and cite only real authority. On a well-supported argument, both are satisfiable and everything goes as expected. But on a fact pattern with no support, thin support, or a novel but rational theory, it can falsify a case to reach the holding rather than tell you it can’t. Give it a case placeholder to fill, and it can unflinchingly falsify a case to fill that placeholder rather than tell you it can’t. The model is mathematically incentivized to produce an answer, and it will randomly drop directives to do so even if that answer is fabricated.
It is simply how the stochastic mathematical model works.
These are real instructions — pulled from firm prompt libraries, CLE handouts, and the “AI policy” memo that went around after the first sanctions headline.
>_Draft the argument section. Only cite real cases — do not hallucinate. Double-check every citation before you answer, give me a link for each one, and if you’re not sure about an authority, say so instead of guessing.
Nothing in that paragraph is wrong to want, but every clause fails for its own reason.
The model has no way to sort the citations it can support from the ones it falsified — both come out of the same probability map. You’ve asked it to filter on something it can’t see about its own answer.
This names the symptom and asks for the symptom to stop. Nothing in the request gives the model access to the authority it was missing, so the gap that produced the invention is still there. Even worse, that's like telling a child "don't lie" — the model will still produce a falsehood, but it will now be more confident about it.
The confidence is written the same way the citation is. A model that could tell you it wasn’t sure would already have known not to invent the case — and it will equally say “I’m confident” about the fabricated one because it has no way of gauging its own confidence.
The check is another pass over the same material. It re-reads its own answer, the mathematical probability of those case names was probable the first time, and it will be again. So it reports back that everything is in order — usually with a fresh summary of the case that doesn’t exist. There's no external verification, so the model is confirming its own work rather than checking against reality.
A link is one more claim in the answer. It can 404, it can open a different case, or it can open the right opinion that simply doesn’t contain the quoted words. It might actually be linking to a complaint rather than a court case, or it might be linking to an article written by a non-lawyer. Until somebody opens it and reads the page, it’s a promise, not proof.
The strongest of the bunch — a real constraint, and it cuts invented case names sharply. What it doesn’t constrain is the quoting: a real case from your memo, a pin cite that’s off, language twisted into something the opinion never says, or dissent cited as the majority.
Some lawyers will try to ask the model to verify its own work — “check each citation before you answer,” or use a follow-up message like “are these cases real?”
What comes back is another probability map answer, not an actual lookup. The same map that produced the citation produces the confirmation, so it confirms. Push harder and it can describe the fabricated case for you — the posture, the holding, a plausible quote — because producing that description is the same kind of task as producing the cite was. Agreement between a model and itself is not evidence. It’s the same claim, twice.
There's been a rush to jump to a technology called MCP because of the convenience of hooking it into the LLM. The hope is that MCP can call a database to look up citations. But hope doesn't replace results.
The deadpool of attorneys sanctioned for legal citation hallucinations is full of those who asked. The question was answered confidently, in writing, by the thing that made it up.
Add the line, run twenty ordinary research questions, and the cites come back clean. That result is real — and it proves almost nothing. On a well-trodden question with abundant authority behind it, the model was likely to cite real cases anyway; the instruction is riding along, not steering.
The brief that draws sanctions is rarely that brief. It’s the novel theory, the unusual jurisdiction, the element with no case squarely on point, the argument that has to reach. Those are the conditions where support runs out — which is exactly where the latent conflict causes falsified citations. So the instruction earns your trust on the easy work and obliterates your trust where it matters most.
A test that always passes where failure is impossible is not measuring anything. The conditions that raise a brief’s hallucination rate are the same conditions that make the instruction fail.
Suppose your citation directive worked 90% of the time. You still couldn’t tell which is fabricated. Fabricated citations tend to obey citation formatting rules: a real-looking case name, a plausible reporter and volume, a page number, a court and a year in the parenthetical.
So now the portion of your citation directive that the model can satisfy is also the part that makes the fabricated citation harder to spot, while the part of the directive it can’t satisfy is the one you were relying on. The answer comes back looking more credible, not less.
And existence is only the loudest and easiest failure to find. A real case cited for a proposition it never reached, a real opinion with a quotation that isn’t in it, the right words attributed to the wrong page (or even the dissent), a statute quoted in a version that was amended years before your filing date — none of those leave a mark on the page either. The directive is a preference, not a proof.
There's plenty of cases published to the internet, why not just ask?
Asking for a source link with every citation sounds like a genuine step up, but that can't close the loop either. The link is no less probabilistically falsifiable than a citation. Or it can open a different case with a similar name. Or the link opens the correct opinion, and the sentence in your brief inside the quotation marks isn’t anywhere in it — tightened, merged from two passages. Or the quote is lifted from a dissent and cited as the majority holding.
That is the step the prompt alone can't ever do for you, no matter how it’s worded. The verification must to happen against the authority itself, and the LLM can't do that.
Some free packages or LLM skills claim to look up citations on one of the major databases. But there's a major misconception here -- the largest cheap/free APIs right now simply check that the volume + reporter + page exists. Nothing more. No case name, no court, no year, no pin cite, no quotes language, and certainly not summarized language, or whether a quote is lifted from dissent rather than the majority. It is a very weak check that can be easily fooled by a fabricated citation -- all that it needs is to land on real combination of volume + reporter + page.
Now in practice, the citation directive does tend to lower the fabrication rate a little. But it also lowers how hard anyone checks a lot. The associate who knows the prompt says “only cite real cases” spot-checks four cites instead of all thirty-four. The partner hears that the firm’s AI policy addresses hallucinations and signs. The line in the prompt slept the trojan horse past the humans.
A control that fails silently while raising everyone’s confidence is worse than no control at all. And when a hallucinated citation is challenged, the conversation is about the document that was filed, not the intent with the prompt that was used to create it.
Benchmarking LLM hallucination rates in legal contexts is something we do regularly as it's part of our research. Hallucination rates vary by a lot of factors, including LLM model, the prompt, and the legal topic. As we stated above though -- these systems are stochastic. They are not deterministic. You can run the same exact prompt 1,000 times and get different results each time. Sometimes the hallucination rate is 0%, sometimes it's 100%, and the rest, it's logically somewhere in between.
Courts and increasingly opposing counsel are looking at every single brief, every single citation. Rough estimates put substantive motions/briefs somewhere between 15,000 to 25,000 filings per day in the US. If even 0.1% are drafted with AI and have a hallucinated citations, that's still 15-25 briefs every day being filed with hallucinated citations.
The consequences are disastrous. The sanctions and fines are already in the millions of dollars, along with suspensions, and even a case of disbarment. And that's before getting to the reputational suicide.
Anthropic makes the Claude AI stack, and yet Anthropic's counsel was found to have filed a brief with a hallucinated citation in Concord Music Group, Inc. v. Anthropic PBC, 5:24-cv-03811, (N.D. Cal.). Are you ready to trust AI alone when they can't even trust themselves?
The fix for stochastic slop cannot ever be more instructions that go through the same stochastic process. That mathematically cannot ever work. The only solution is verification that happens after the draft exists, against the authorities themselves. Verbatim reads a finished brief and reports, for every authority cited, determines whether the cite is real and whether the quoted language actually appears at the pin cite — full cites, short forms, Id., supra, statutes and regulations at the version in effect on your filing date. Every verified cite carries a link straight into the source, so the page you were trusting is one click away.
So keep the instruction. Or don't. But stop treating it as the thing standing between your firm and a show-cause order, and put a real check where that belief was.
Verbatim reads a finished brief and reports, for every authority it cites, whether the cite is real and whether the quoted language actually appears at the pin cite — so a fabrication surfaces on your screen, not in a show-cause order. Bring a brief and we’ll walk you through the report.