PAPER A · AI IN PRODUCTION

Why 88% of AI pilots die — and the 12% protocol

KET · 04 — INSIGHTS

The meeting where it dies

The meeting where an AI project dies is never called "the cancellation meeting."

It's called a "progress review." Eight people in the room, the demo happened four months ago, everyone had applauded. This time someone from IT asks a question that isn't hostile, just precise: "when the agent sends an email to a customer, where do we find a record of it?"

Silence. Then compliance follows up: "and your test data — where did it come from?"

Nobody says no. Nobody says yes either. The project moves to "needs consolidation," the sponsor rolls off to another topic next quarter, and six months later all that's left is a Git repository nobody opens.

I saw this scene long before AI. At Amadeus, at Air France-KLM, at Enedis, it's the same meetings, the same questions, the same silences — on subjects that had nothing to do with language models. What changed with generative AI is how fast you get to the demo. What hasn't changed at all is what you have to prove to go any further.

The number, and what it doesn't say

88% of AI pilots never reach production. The figure is everywhere right now, it's solid, and the reasons teams give themselves are always the same three: evaluation, governance, reliability. Put plainly: we can't measure whether it works, we don't know who's accountable for what, and we don't know what happens when it gets things wrong.

Two clarifications before going further, because this number gets misused.

One. Some of those 88% should die. Plenty of pilots are launched to learn, or to answer an internal push — "do something with AI." A pilot that proves a use case isn't worth it has done its job. The waste isn't a pilot stopping; it's a pilot that worked stopping for reasons that could have been handled on day one.

Two. The inverse number is more interesting. Organisations that make it across measure returns that look like nothing else in an IT budget. That's not an argument to rush — it's an argument not to let the 12% that deserve it die.

So the real question isn't "how do I save every pilot." It's: is mine dying because it was worthless, or because nobody prepared it to scale?

What separates a demo from a system

A demo convinces on twenty hand-picked cases. Production faces thousands of real ones, every day, including the ugly ones: the customer who replies with an attachment, the crooked scanned PDF, the duplicate invoice, the user who asks something off-topic, the service that goes down on a Friday at 6 p.m.

Between the two, there are six workstreams. None are glamorous. None show up in a proof-of-concept budget. And they're exactly the six the project gets rejected on.

1. Error handling

The question to ask early: what happens when the agent gets it wrong? Not "if." When.

You need a written answer to three sub-questions: how you detect it, who's alerted, who fixes it. In most pilots I look at, the answer is: nobody detects it, because there's no visible difference between a correct answer and a confidently phrased wrong one. That's the most dangerous property of these systems: they fail while keeping the same tone.

2. Traceability

Knowing what the agent did, on which data, and why. Timestamped, viewable, without having to call a developer.

This is the workstream that unlocks all the others. An IT department won't authorise what it can't observe, and a regulator won't accept what can't be reconstructed. We have a name for it — the proof ledger — and it's the first thing we wire in, before even improving the model's performance.

3. Human validation

In the right places. Not everywhere, not nowhere.

Everywhere is the "let's put a human in the loop" trap that reassures everyone in the meeting and turns automation into extra work: after three weeks, the person is rubber-stamping without reading. Nowhere is the compliance nightmare. The work is identifying the two or three points where a decision is costly or irreversible, and putting the control there — only there.

4. Access security

The agent touches your real systems: ERP, CRM, email, customer database. With which permissions, granted by whom, revocable how?

Plenty of pilots run on the personal credentials of whoever built it. That's fine in a pilot. It never survives a security review — and rightly so.

5. Cost control

An agent without guardrails spends without limit. A loop left open, an over-long document, a usage spike, and the monthly invoice has nothing to do with the business case any more.

It's not only about money: it's about credibility. A project whose unit cost nobody can forecast won't survive a budget arbitration, whatever its performance.

6. Documentation

The kind IT, the auditor and — now — the regulator will ask for. Written to be read by an outsider, not to fill a shelf.

It's the most postponed workstream and the most expensive to catch up on. Reconstructing, after the fact, the architecture choices, the data used and the reasons for a decision made eight months earlier is archaeology. Done as you go, it's an hour a week.

The 12% protocol

What I've just described reads like a shopping list. It's actually always the same mechanics, in the same order. We apply it to AI systems the same way we apply it to regulatory obligations, because it's the same muscle.

Inventory — what actually exists? The systems in use, the components, the data flows, the access, the owners. This step always surprises: in a mid-sized organisation, there's systematically more AI in service than IT knows about. Teams have plugged in tools, vendors have shipped some, a subscription is lingering somewhere. You don't govern what you haven't mapped.

Process — who does what, when, with which validation? Error handling, checkpoints, escalation, on-call. With one non-negotiable rule: a process that only lives in a PDF is dead. It has to execute — a real alert, a real ticket, a real person at the end of it.

Documentation — what do you show an auditor? The technical file, the justification of choices, the update policy. Versioned, dated.

Proof — can the system demonstrate itself that it works? Logs, decision traceability, alerts, periodic reports. This is what turns "trust us" into "verify it yourself." Declared compliance wears out. Demonstrated compliance holds.

The order matters. You often see teams start with documentation, because that's what they were asked for — and produce a handsome file describing a system nobody actually inventoried. It doesn't hold up for ten minutes in front of someone asking the right questions.

The honest timeline

An audit of this kind takes two to three weeks. It doesn't produce a slide deck: it produces the current state, the list of gaps, and a costed plan. The productionisation that follows typically takes six to ten weeks depending on scope.

And the other half has to be said too: sometimes the audit concludes the pilot doesn't deserve production. A use case too narrow, input data too dirty, a real gain well below the cost of running it. That's a useful conclusion, and it beats another six months.

The real starting point

If you have a pilot waiting, there's one question to ask before all the others, and it isn't technical:

who, internally, will own this system once it's running?

Not the sponsor. Not the vendor. The person who'll get the alert at 6 p.m. on a Friday. If that name doesn't exist, the project will die in production rather than before it — that's the only difference.

The six workstreams can be handled. The four steps can be followed. But a system without an owner dies too, and no protocol fixes that.

Read nextAI Act & CRA: what's due in 2026, what slipped to 2027: the regulatory calendar this same protocol is built to meet.


Keteris takes AI pilots into production: readiness audit, build, operations. We apply this method to our own products — Le Reglo, our regulatory-watch platform, has been running in production since 2024.

A pilot waiting? Book a 30-minute scoping call — answer within 24 business hours, straight opinion included.

All insights