Why most AI pilots never reach production

Pilots rarely fail on the model. They stall on ownership, workflow and trust. A teardown of why, and what gets one into daily use.

Most AI pilots don’t fail. They fade. The demo impresses the room. The pilot runs for a few weeks. Someone writes up the results. Then everyone goes back to working exactly as before. Nobody decides to kill it. It just never becomes part of anyone’s job. It’s the pattern Momentem was set up to break, which is why we build the system and then embed engineers with the team until it runs in production. This is a teardown of why pilots stall, and what the ones that make it do differently.

The model is rarely the problem

When a pilot stalls, the model usually takes the blame. It wasn’t accurate enough. It made something up. It wasn’t ready. Sometimes that’s fair, but in my experience it’s rarely the real reason. In the demo, the model mostly did what it was asked.

What failed was everything around it: the problem that was picked, who owned it, how the work changed, how anyone knew it was working, where it ran and whether people trusted it.

Six places pilots stall

It started with a tool

Someone saw a demo, the company bought access and a team went looking for somewhere to use it. Because the pilot was defined by the technology, success meant “it works”, not “this job got better”. A tool looking for a problem can always find a demo. It rarely finds a budget for year two.

Nobody owned it

It belonged to an innovation team, the IT department or a vendor. The people whose work was meant to change were consulted, maybe, but they weren’t accountable for the result. When the pilot ended, keeping it running was nobody’s job, and nobody involved had the authority to change how the work gets done.

The work didn’t change

The AI was bolted onto the old process as an extra step. People did their job the way they always had, then checked the AI’s output on top. That’s more work, not less. If a pilot doesn’t remove a step from someone’s day, it adds one, and people quietly route around anything that adds work.

Nobody could say if it was working

There was no agreed definition of good before the build started, so the pilot was judged on anecdotes: a few impressive examples, a few embarrassing ones. The embarrassing ones travel further. Without a baseline and one number to watch, you can’t show that version two beats version one, and you can’t make the case to whoever has to sign it off.

It lived outside the real systems

The pilot ran in a separate tab, on exported data, in a sandbox. Getting it into the inbox, the CRM or the ticketing system, with the right permissions, logging and security review, is a large part of the job. That part often wasn’t in the plan or the budget, so the pilot succeeded on paper and then hit a wall.

People didn’t trust it

The people expected to use it had seen it be confidently wrong, with no easy way to check or correct it. Some also wondered, fairly, whether it was there to replace them. Trust doesn’t arrive with a launch email. It builds when people can see what the system did, stop it when it’s wrong and watch it get things right over time.

What gets a pilot into daily use

The pilots I’ve seen make it into daily use were less ambitious and far more structured. They tend to share five things.

“A pilot proves the model can do the work. Production proves the team will use it.”

One workflow

Not “customer service” or “marketing”. One job that happens often and hurts: the first reply to a refund request, the weekly performance report, sorting inbound leads. Narrow enough to finish, frequent enough to matter and easy to measure.

One owner

A named person who does or runs that work every day, not the executive who approved the budget. They decide what good looks like. They can change the process. They’re the first to notice if it stops.

One number

Agree it before anyone builds anything, and measure where it stands today. Time to first reply. Hours spent each week. Errors in a sample of the output. Pick a number the owner already cares about, and review it on the same day every week.

A human approval step

Let the AI do the work, and have a person approve anything that goes out. This gets you live sooner, because the system doesn’t have to be perfect on day one. It only has to be easy to review. Every approval and every edit also shows you where it’s improving and where it isn’t. I’ve written separately about how to design that step so it stays fast.

Someone embedded until it runs

This is the one most pilots skip. An engineer works alongside the team doing the job, watches where the system breaks and fixes it in days, not at the next quarterly review. They stay until it runs without them, not until the pilot report is written. It’s how we work at Momentem, but you don’t need us to borrow the idea. You need someone who owns the outcome and stays close to the work.

Before you start the next one

Answer these on one page before anyone opens a tool:

  1. What is the workflow, in one sentence?
  2. Who owns it, and is it their work that changes?
  3. What’s the number, and what is it today?
  4. Which step does this remove from someone’s day?
  5. Which system will it run inside?
  6. Who approves the output, and how long should a review take?
  7. Who stays with the team after launch, and until when?
  8. What result would make you stop?

If you can’t answer the first three, you’re not ready to build. You’re ready to talk to the people who do the work.

And when you think you’re finished, try one last test. If you switched it off tomorrow, would anyone complain? If not, it’s still a pilot.

If you’re planning a pilot now

If you’d like a second pair of eyes on your plan, book a call and bring the one-page version with you. It’s thirty minutes, and you’ll leave with a clear next step whether or not we work together.

Written by MuddaaFounder of Gridwork, Momentem, Leila and Onshow. 11+ years in marketing, tech and AI.
Get the next one in your inbox.One email a week. Notes from the build, no spam.