Why Most AI Pilots Never Make It to Production
Share:

A few months ago I had a call with a founder whose AI project had stalled. Everything in the pipeline worked, from data ingestion to model output and the demo actually looked promising. The problem was clearly bridging the gap between the demo and a production system. He was, understandably, pretty frustrated; he talked about a developer he'd hired the year before, paid a decent chunk of money, and who had, as far as anyone could tell, sent zero emails in the outreach system he was supposedly building. Nothing badly written, nothing half working, just nothing at all going out the door for weeks, until he finally started asking why the pipeline looked so quiet and got an answer he didn't like.
By the time he got to me, he'd been burned enough times that he opened the call almost daring me to disagree with him. Something along the lines of, I've spent real money trying to figure out if this stuff actually works, and at this point I'm starting to think the answer is no.
I've had a version of that same conversation three separate times this year, in three different industries, and I recognize the posture now before the person finishes their first sentence. Once you get past the frustration and sit with what actually happened, the model is almost never the reason things fell apart. Sometimes this surprises people (increasingly less so recently with the speed at which models are advancing), but after years building AI teams and shipping AI systems, I've noticed that most AI pilots don't fail because the models aren't capable enough. They fail because the surrounding engineering and delivery process isn't designed for production.
I've spent close to fifteen years around this particular failure mode. Screening AI talent at Toptal, building out an entire department of it at Turing for clients who needed people who actually knew what they were doing, and now running delivery here at Eventum. You start to notice a pattern after you've watched it enough times that it stops feeling surprising and starts feeling almost routine. There's a small set of places where AI pilots quietly die, and none of them need a worse model to kill the thing. They mostly just need nobody to have done a fairly unglamorous piece of homework early enough in the process.
I figured it was worth writing down, partly for whoever reads this before greenlighting their next one, and partly because writing it down is how I keep myself honest about it too.
Nobody wrote down what "done" meant
This is the one I see most often, and it rarely feels like a mistake while it's happening. A team starts building toward something that mostly exists in one person's head, usually whoever's the most excited and most senior in the room. Everyone else is inferring the spec from Slack messages, half finished sentences in stand-ups, and a general sense of the direction. For a while that's fine. Then, months in, the system does something genuinely impressive on a Tuesday demo, and the room goes quiet in a way that isn't good, because it turns out that wasn't actually what anyone needed. It was just the thing that happened to be easiest to show off that week.
I watched a project grow from a serious scope commitment into something several times larger than what was originally priced, purely because nobody had a signed amendment defining what "bigger" even meant along the way. There was just a shared, unspoken sense that better kept moving a little further off every time someone got close to it. Every demo cycle reset the bar instead of clearing it, which wears people down in a way that's hard to explain unless you've lived through it.
What we changed after watching that happen is almost embarrassingly simple, honestly. A written spec, with acceptance criteria you could actually test against, exists before a single engineer gets staffed. Not a vision document with nice language in it. Just a plain list of things that are either true or not true on the day someone checks.
The person who controls the budget never actually saw a demo
I've sat in plenty of discovery calls with people who are smart, technically sharp, and genuinely excited about a project. An engineering lead, a chief of staff, someone who champions the whole thing internally with real conviction. And more often than I'd like, it turns out they're not actually the person who can say yes. Somewhere above them sits a board, or a CFO, or a founder who controls the money and has never once joined a call.
We lost a deal that both sides genuinely believed was close to locked. In writing, even, something close to "you're our top choice." The person who actually controlled the budget never showed up to a single conversation, and by the time internal politics got involved on their side, we had no relationship there to draw on. Nobody misled us, to be clear. The client didn't lie about anything. We just never asked the slightly awkward, fairly obvious question early enough: who actually signs this, and can we get fifteen minutes with them before we go deep into scoping.
I ask that question on the first call now, not the fifth. If someone can't name who signs the check, that tells you exactly what the next call needs to be about.
Scope grew, and nobody ever re-signed anything
This one is almost always well intentioned, which is part of why it's easy to miss while it's happening. The client sees early progress, gets excited, asks for just one more thing. Nobody wants to be the person who tells an excited client no, so the team says sure, we can add that. Six or seven "just one more things" later, the project looks nothing like what was priced or timelined or staffed, and not one of those additions ever got written down anywhere.
I don't really think of this as a communication problem. The conversations themselves were fine. What was missing was a five minute habit: turning "sure, we can add that" into an actual line that changes either the price or the date before the work on it starts. Skip that habit enough times and a project quietly rots from the inside while everyone's still smiling on the calls.
A demo going well got mistaken for the thing actually working
This might be the sentence that costs the most money in this business, said in some form after almost every stalled pilot I've looked at: it seemed good in the demo. A curated set of inputs, a rehearsed flow, a feature quietly switched off at the last minute because it wasn't quite ready and nobody wanted to say so out loud. Everyone in the room, sometimes including the people who built it, starts believing the system is closer to finished than it actually is.
I've watched a team ship a demo with something disabled rather than admit it wasn't working yet, and watched that cost them the client's trust permanently once it came out, which it always does eventually. Saying "this part isn't working yet, here's what's left" in the room turns out to be a lot cheaper than the conversation that happens two weeks later, once everyone finds out anyway.
These days we build something we call a truth table before every demo. What's real, what's mocked, what's disabled, shared with the client directly rather than just discussed internally beforehand. Evaluation means offline tests plus actual production monitoring over time. It doesn't mean a demo went well in front of an audience that wanted to be impressed.
There's a fifth one, closer to home for me
Briefly, because it's the piece closest to where I actually spend most of my time. Sometimes a pilot dies because the team assembled for it was built to produce a good demo, not something that survives after the demo ends. I've screened enough AI engineers to usually tell, fairly early into a conversation, the difference between someone who talks fluently about retrieval architectures and someone who can explain, plainly, what happens when the underlying data drifts six months down the line. The first kind builds you an impressive pilot. The second kind builds you something that holds up in a real, messy production environment. Most teams don't find out which one they hired until it's already expensive to find out.
What I'd actually check before greenlighting something
If I had to boil this down, it's five questions I run through now, more habit than ceremony. Is "done" written somewhere as testable criteria, or does it live mostly as a feeling in someone's head. Has the person who controls the budget actually seen a live demo, or only heard about it secondhand. Is there a real process for re-scoping midway through, or does scope just quietly grow. And does the team know, in writing, the difference between demo-ready and production-ready, or is "it's basically working" doing a lot of the heavy lifting.

If more than one of those feels shaky, you're probably not really running a pilot anymore. You're running an expensive improvisation that happens to have AI in the name.
None of this is theoretical for us, which is exactly why I'm comfortable writing it out plainly. We've seen these failure modes firsthand; in projects we've inherited, teams we've advised, and lessons that have shaped how we run our own engagements. A written spec before anyone's staffed. The approval chain mapped on the first call, not the fifth. A short written amendment for every scope change, no matter how small it feels at the moment. A truth table before every demo, shared honestly with the people paying for it. It's not a more exciting way to run a project, if I'm honest. It's a way of finding out on day five whether something is real, instead of finding out on month five, once the budget and the trust are both already gone.
If you're staring at a pilot that's stalled, or you're about to greenlight a new one and want someone to pressure test it first, we run a scoped, priced discovery pass built around roughly these four questions before you commit real budget to a build. You can reach out at eventum.ai. I'd genuinely rather tell you honestly that something isn't ready yet than let you find that out the expensive way, again.

