AI Errors Have Reached 1,395 US Court Cases. The Problem Is No Longer Just Hallucination

Illustration of a lawyer reviewing legal documents beside an AI-enabled laptop in a US courtroom.

AI-generated errors have now been publicly identified in at least 1,395 US court cases.

That number does not come from the judiciary. It comes from an independent tracker maintained by researcher Damien Charlotin, meaning it is neither an official count nor necessarily a complete one.

But the cases it captures point to a problem the legal profession has had three years to understand.

Generative AI can invent things.

The first major sanctions over fabricated AI-generated legal citations arrived in 2023. Since then, courts, bar associations and law firms have repeatedly warned lawyers that using AI does not transfer their professional responsibility to verify what they submit.

Yet fabricated cases, inaccurate quotations and invented evidence are still reaching courts.

Three incidents this month show how widely the problem can spread.

In New Mexico, a defence lawyer was rebuked after material submitted in a murder appeal included fabricated information about what police witnesses had supposedly said.

In Oklahoma, a judge acknowledged using AI-assisted research while preparing an order that contained citations to cases that did not exist.

And in California, a lawyer representing State Farm was fined after filing material containing inaccurate citations and quotations.

These are different failures involving different roles inside the justice system.

That matters because the problem can no longer be reduced to inexperienced lawyers discovering ChatGPT and trusting whatever appears on the screen.

AI-generated errors have surfaced in work produced by lawyers, prosecutors, experts and judges themselves.

The technology is getting better.

The verification problem has not disappeared with it.

Modern AI systems can search documents, summarise enormous case files, compare arguments, draft legal language and locate potentially relevant authorities far faster than a person working manually.

Those capabilities create an obvious attraction in a profession where time is expensive.

A lawyer who can reduce several hours of preliminary research to minutes can spend more time on strategy, clients or other cases. Courts facing large caseloads have similar incentives to use tools that accelerate research and drafting.

But the productivity gain depends on what happens next.

If AI performs an hour of research in five minutes and a lawyer then checks the authorities, quotations and factual claims against the underlying sources, the technology may genuinely have saved time.

If the verification step disappears because the output looks convincing, the calculation changes.

The system has not eliminated the work.

It has transferred some of the risk created by not doing it.

In law, that risk can land on somebody else.

A fabricated citation can waste a judge’s time. An invented quotation can distort an argument. False information about evidence can affect how a court understands a case. At the extreme, an error can enter proceedings involving somebody’s liberty, finances or legal rights.

That is why legal hallucinations matter differently from an AI making a mistake while answering an ordinary consumer question.

The output enters a system in which claims are supposed to be supported by evidence and authorities.

Lawyers already have obligations designed for precisely that environment.

They are responsible for the documents they file. Using software to prepare those documents does not make the software responsible for their accuracy.

The same principle existed long before generative AI.

A lawyer could not defend an inaccurate filing by explaining that a junior employee made the mistake, that a database returned the wrong result or that somebody copied information incorrectly. Tools and assistants can contribute to legal work without inheriting the professional responsibility of the person submitting it.

AI changes the scale of that problem because it can produce plausible falsehoods extraordinarily quickly.

A nonexistent case can arrive with a convincing case name, citation, procedural history and quotation. Unless somebody checks the underlying authority, the fabrication may look very much like legitimate legal research.

That makes fluency part of the risk.

Bad information that looks obviously bad is relatively easy to catch. Bad information written in the language and structure expected from a legal professional is more dangerous precisely because it does not immediately announce itself as unreliable.

The growing number of publicly identified cases therefore tells us something beyond the reliability of AI models.

It tells us about the systems being built around them.

A safe legal workflow does not have to assume AI will never hallucinate. It has to assume that it might, and place verification between machine-generated research and anything submitted to a court.

That could mean requiring every cited case to be opened in an authoritative legal database. Quotations can be checked against the original judgment. Factual assertions can be traced back to the evidential record. Firms can require lawyers to disclose internally when generative AI has contributed to research or drafting so that appropriate review takes place.

None of those controls eliminates the productivity benefits of AI.

They determine whether those benefits are real.

That distinction matters as businesses increasingly describe AI in terms of hours saved and work automated.

Time removed from the first stage of a task is not necessarily time saved across the entire process.

If reliable use requires checking the machine’s work, verification belongs inside the productivity calculation.

The legal profession makes that unusually visible because its mistakes eventually appear before judges.

Other industries may have the same workflow problem without producing a public court order every time something goes wrong.

An accountant relying on an invented regulation, a researcher accepting a fabricated source or an analyst incorporating a false statistic faces the same underlying issue: automation can accelerate production faster than it accelerates verification.

That does not make the productivity gain imaginary.

It changes where the human work belongs.

The most effective use of AI may not be replacing professional judgement but moving it. Less time can be spent producing a first draft or finding possible answers, while human attention moves towards checking sources, resolving ambiguity and deciding what can safely be relied upon.

The 1,395 identified court cases do not establish how often AI is used successfully in American law. There is no comparable public count of the filings in which lawyers used AI, checked its work and submitted something completely accurate.

So the tracker cannot tell us an AI error rate.

What it does show is persistence.

Three years after courts began imposing sanctions for fabricated AI authorities, the profession is still producing examples in which output from systems known to invent information reaches formal legal proceedings.

At that point, hallucination is only the first failure.

The second happens when nobody catches it.

AI can produce work faster than professionals could produce it themselves.

The question for law, and increasingly for every profession adopting the technology, is whether verification can keep up.

Sources

Share this story