Skip to content
German Orlov
All writing

Essay · 4 August 2026

Why agentic AI pilots stall in regulated enterprises

The pilot is not the hard part. Four things decide whether an agent reaches production, and none of them is the model.

Most agentic AI pilots in a regulated enterprise succeed. That is the problem.

A pilot proves a model can do a task. It says nothing about what decides whether the system reaches production. I have watched each of the four below stop work that deserved to ship.

The pilot answers a question nobody was asking

The question inside a pilot is "can it work". The question in production is "who is accountable when it does not".

Every enterprise I have worked in can answer the first in six weeks. The second takes longer, and it is not a technical question.

Four things that have to be true

The use case has an owner and a baseline. Not a sponsor — an owner, with a number they are already measured on. If nobody knows the metric before the agent, nobody can say what the agent changed. Our voicebots went into a live acquisition funnel with a known conversion baseline. That is the only reason the lift means anything.

The data is reachable, not merely present. Regulated groups have the data. It sits behind perimeters, residency rules, and three systems that disagree. An agent that needs a human to fetch its context is a demo with extra steps. Most of the work I do before an agent is integration work.

Someone has decided what the agent may get wrong. This is the one that stalls pilots quietly. A model with no defined failure behaviour cannot be signed off, so it waits. Decide the handoff, the confidence threshold, and who reads the logs. Then legal has something to approve rather than a capability to fear.

The team that will run it was in the room before it was built. Adoption is not a launch activity. I shadowed operations teams before designing workflows they now use daily. That is slower at the start, and it is the only version that survives the first month.

What this means for the next pilot

Stop running pilots designed to succeed. Run the smallest thing that touches the real perimeter — real data, real users, a real approval. It will look less impressive. It will also tell you whether you have a system or a slide.

The gap between a pilot and production is rarely the model. It is ownership, integration, defined failure, and the people who inherit it. Those are the four I assess before an engagement starts, not after.

Have a problem worth solving at scale?

I work with teams turning AI and data strategy into systems that ship.

Start a conversation