Skip to main content
Back to insights
Pedro NogueiraSEP 15, 20269 min read

The AI Maturity Model for Founder-Led Companies

Share:

The AI Maturity Model for Founder-Led Companies

Founders consistently overrate where their company sits on AI. Not out of dishonesty. The only benchmark most of them have is a competitor's press release or a conference demo, and neither tells you anything about your own operation. The stage you're actually in is the only thing that should decide what you do next, and most people never sit down and figure out which one that is.

I use five stages when I'm sizing up a company. Each has a symptom, a mistake that's specific to it, and one thing that actually moves you to the next one.

Stage 1: nobody's decided anything

People across the company are already using ChatGPT, Claude, or Gemini on their own. There's no company position on it. No approved tool, no policy, nobody whose job it is to care.

A law firm I advised was a textbook case. The partner was proud his people were "already using AI to move faster," and they were, in the sense that everyone had picked their own tool based on whatever they'd personally gotten comfortable with. Some on Claude, some on Gemini, no consistency. The part that should have worried him more than it did: they were running client documents through these tools with zero policy and zero logging. My pitch wasn't "you should adopt AI". They'd already adopted it, informally, per person. The decision sitting in front of him was whether to govern what was already happening.

Ask "who owns AI here" and if you get a shrug, or three different names, you're in this stage.

Companies here tend to do one of two things wrong: ignore it completely, or overcorrect with a splashy top-down initiative before anyone's looked at what's already happening on the ground. Both skip the actual first step, which is just finding out what your team is doing with these tools right now, because they already are.

What moves you forward isn't a hire. It's a short look at where AI is already touching the business, sanctioned or not, and naming one person who owns deciding what's worth formalizing.

Stage 2: pilots without a way to judge them

More than one team is running its own AI experiment. Nobody's comparing notes, and each team judges its own pilot by its own standard, which in practice usually means "the demo looked good".

The failure I see most at this stage isn't a bad pilot. It's a pilot nobody priced out before deciding to build on it. A consumer app I worked with had already done their own math and found their planned AI features would cost $35-55 per customer per month to run, against price tiers of $0-20. The normal move at that point is to keep pushing forward on the AI version and hope the cost comes down later. They did the harder, better thing instead: went back and asked where a deterministic classifier or a cheap model would do the same job for a fraction of the cost, and where AI genuinely earned its keep. That's what Stage 2 is supposed to look like, and most companies don't get there because nobody's running the cost side of the evaluation, only the "does it work" side.

If you can't get a straight answer to "how do we know this is actually good, including what it costs us to run", that's the tell.

What unlocks Stage 3 isn't more pilots. It's one shared way of judging all of them, cost included, and someone with the standing to kill one.

Stage 3: the pilot everyone's proud of that won't ship

This is different from Stage 2 in one specific way: a pilot here has already earned internal confidence and buy-in, and it's still not in production months later.

A founder told me flatly he was done with agentic AI. He'd put well into six figures across several outside developers trying to get something working. The system itself, an outreach agent, had been sending messages continuously for weeks and had booked exactly zero meetings.

This isn't exactly a case of "the pilot everyone's proud of that won't ship". This one actually did ship; the problem was more along the lines of "We put something into production without the evaluations/monitoring necessary to know whether it's accomplishing anything”.

This founder was ready to write off the entire category as vaporware, but the model wasn't the problem. Quite simply, nobody had built any way to notice the thing wasn't working, so it just kept running, and motion got mistaken for progress because there was no measurement telling anyone otherwise.

If the same pilot has been "two weeks from launch" for longer than two months, this is where you’re at, whatever the roadmap says.

The mistake at this stage is throwing more prompting at what's actually a systems gap: validation, escalation paths, handling the cases that don't fit the pattern. No amount of prompt tuning replaces an evaluation harness that doesn't exist.

Getting past this means treating the pilot-to-production gap as its own scoped project instead of a finishing touch, and this is usually where an outside person who's made this specific jump before is worth paying for. Not because your team is bad. Because they haven't done this particular transition, and someone who has knows which 20% of edge cases actually matter and which can wait.

Stage 4: One thing works, and only one person understands why

An AI workflow is live, trusted, and depended on daily. It's also usually bespoke, held together by one person's judgment calls, and hard to copy for the next use case.

A property-operations client of ours is a good example of this done well: a confidence-gated agent triaging tickets across their stack, with a human in the loop from day one and explicit testing of where it fails. Anything touching billing or legal goes straight to a person, no exceptions. It works and people trust it. What's still open at this stage isn't whether the system works, it's whether the reasoning behind it (the confidence thresholds, the escalation rules, the specific failure modes they tested for) exists anywhere outside the one build it was made for.

If the person who built it left tomorrow, would it keep running the way it does now? If you hesitate, that's your answer.

The mistake here is assuming one working system proves the company "does AI" now, and then trying to repeat that win ad hoc for the next five ideas instead of pulling out what actually made it work and reusing that.

Stage 5: It's infrastructure, not a project list

AI is an ongoing capability with shared plumbing: evaluation, data access, monitoring, that any new use case plugs into rather than rebuilding from zero.

You'll know you're here because a new idea can go from concept to pilot in a couple of weeks, not because you decided to move faster, but because the scaffolding already exists to support it.

Almost no founder-led company actually needs to get here, and I'd tell you that directly if you asked me. Past a certain size, chasing Stage 5 is over-engineering a capability three future hires don't need yet. The honest answer for most companies I look at isn't "get to Stage 5," it's "you're at Stage 2, stop trying to act like you're at Stage 4." Knowing you don't need the next stage is worth more than the stage itself.

A rougher gut-check than a scorecard

Skip the scoring math. Ask yourself these straight:

  1. If I asked "who owns AI here," would I get one name or a shrug?
  2. Do we know what any of our AI features cost us to run, per customer, at real usage, not demo usage?
  3. Has anything been "almost ready" for more than two months?
  4. If the one person who understands our AI system left tomorrow, would it keep working?
  5. Have we ever actually killed a pilot because the evidence said to, or has every one either shipped eventually or just quietly died?

A shrug on #1 and no answer on #2 means you're early, Stage 1 or 2, and that's fine, most companies are. A "yes" on #3 puts you in purgatory no matter what else is true. A confident "yes" on #4 means you're further along than most and the real question becomes whether that one system can turn into a pattern.

Where this actually goes

If you're early, the useful move isn't hiring anyone. It's getting an honest outside look at where you stand, which is what an AI Opportunity Audit is built for, and a legitimate outcome of that audit is telling you not to build anything yet.

If you've got something working but nobody senior enough to turn it into a pattern instead of a one-off, that's the point where a fractional AI CTO retainer actually pays for itself, not for writing code, but for making the calls on what to formalize, what to kill, and what's next.

Most of the money I've watched founders waste on AI didn't go to a bad model. It went to skipping the stage they were actually in.

Summarize with AI:

ChatGPTGrokGeminiClaude
Related service

Discovery, architecture, build, evaluation, deployment, handoff. Senior technical ownership end-to-end.

Diagonal halftone representing the AI project delivery flow.