← Blogs
Problem

Why 95% of enterprise AI pilots fail inside regulated institutions

Everyone blames the models. The evidence points somewhere else: a structural mismatch between how generic AI behaves and what regulated work actually requires.

Rows of financial documents and printed reports on a desk

There is a number worth sitting with. In its 2025 study The GenAI Divide, MIT reviewed 300 public enterprise AI deployments and found that 95% delivered no measurable business impact. Only one in twenty produced real, attributable value.

The instinctive read is that the models were not good enough. The study says otherwise. The determining factor was not model quality, it was implementation. Generic tools bolted onto the side of real work, impressive in a demo, and quietly abandoned in production within months.

Inside regulated enterprises specifically, banks, fintechs, and payment processors, a second and harder problem sits underneath that 95% figure, and it gets far less attention than it should.

Generic AI is built to always answer

A general-purpose model is built to always produce something. Ask it a question outside its knowledge and it will not stop to say "this one is unknown." It fills the gap with an answer that sounds plausible.

In most consumer contexts, that trade is acceptable. A slightly wrong answer is a minor inconvenience. Put the same behaviour inside a regulated institution, and the calculus changes entirely.

A confident wrong answer about a regulatory deadline is not a typo. It is a fine. A misread threshold on a suspicious-transaction alert is not an inconvenience. It is a filing that should have happened and did not.

The work these institutions run on has a specific character that makes guessing intolerable:

This is high-stakes, low-judgment work. It follows published rules the overwhelming majority of the time, and it is still done by hand, at significant cost, because errors are expensive and a general model cannot be trusted to stay inside the boundaries of what it actually knows.

The question changes once liability enters the room

Most of the industry is still asking "can AI do this work?" Technically, the answer is usually yes. Current models are capable of reading a document and drafting a response.

That is not the question the person who has to approve the output is asking. The compliance officer, the risk lead, the person whose name is attached to the outcome, is asking something narrower and harder to satisfy:

"Can this be done in a way I am willing to sign?"

Those are two different questions, and the gap between them explains most of the 95% figure. It is why an impressive demo rarely becomes the system running production casework on a Monday morning. The demo answers the first question. Almost nothing in the market currently answers the second.

The uncomfortable part

If the problem were simply "the models need to get smarter," it would resolve itself on a normal industry timeline. Every frontier lab is already working on exactly that, and capability does improve every few months.

That is not the constraint. A more capable model that still guesses under uncertainty is still a model that guesses. Increased fluency does not make a risk team more willing to trust a filing to it. If anything, a more confident wrong answer is more dangerous, because it is harder to catch before it causes damage.

Which leaves the question that actually matters: if the fix is not a smarter model, what is it? What has to be true structurally, not how clever the system is, but how it is built to behave, before "interesting demo" becomes "we can run this in production"?

That question has a concrete answer, and it has very little to do with the metrics AI headlines usually lead with. It concerns what the system is engineered not to do: not decide the things it should not decide, show its reasoning at every step, operate under limits a human sets and can lock permanently, and stop rather than guess when it encounters something it was not built for. See what AI has to be before a compliance officer signs off for that list in detail.

Before restraint, though, there is a more basic filter regulated buyers apply, one that has nothing to do with output quality at all: where does the data actually go. That question, and why it quietly ends more AI projects than the technology ever does, is covered in the question that kills AI projects in a bank.

Source. MIT NANDA, The GenAI Divide: State of AI in Business 2025, based on a review of 300 public enterprise AI deployments.

CaseClear is built around the structural answer to the question this piece raises: rules decide, AI drafts, a human approves, and every action is logged and reproducible.

See it run against your rulebook

Point to one module, share 20 to 30 anonymised cases, and see it run against your rules.

Join waitlist →