← Playbook

Eight of ten AI pilots die at month three. Here is the autopsy on all of them.

They do not fail for technical reasons. They fail in five specific, predictable, entirely preventable ways, and you can check for every one of them before you sign anything.

2026-08-09/7 min read/Levelbrook AI Practice

The month-three death is so consistent it is almost a schedule. Month one is enthusiasm. Month two is a demo everyone claps at. Month three is when someone asks what it has actually saved, and the honest answer is nothing, and the project quietly stops being mentioned in the update.

Here are the five causes of death, in order of how often we find them when we get called in afterwards.

Cause 1: read-only from the start (the biggest killer)

The AI could answer questions. It could not do anything. So a human still opened every ticket, read the AI's suggestion, and then did the work by hand anyway.

This adds a step. It has negative ROI by construction, and everyone can feel it long before anyone can prove it. The reason it happens is understandable — write access felt risky, so it got deferred to "phase two" — but deferring the write is deferring the entire value.

The check: before signing, ask what action the system will take in production, in week four, without a human doing the clicking. If the answer is a variant of "it suggests", you are buying phase one of a project whose value is entirely in phase two.

Cause 2: no baseline, so no possible proof

Nobody measured the before. Now it is month three, someone wants the ROI, and the only available evidence is that people feel it is faster. Feelings do not survive a budget review.

This is the most avoidable death on the list and it costs one week of counting. Median handling time per action type, volume per week, cost per resolution, first response time, escalation rate. Take it before anyone installs anything.

The check: what numbers are we taking, this week, before the build starts, and what will they be compared against?

Cause 3: it was scoped as "AI for support" instead of an action

Broad scope means no completion criterion, so it is never done, so it is never proven, so it dies of exhaustion rather than failure. "Automate support" is not a project. "Handle tracking enquiries end to end, at level 3, with under 0.5% recall" is a project — it can be finished, measured, and pointed at.

The check: can you write down the exact success condition for automation number one, with a number in it, before starting?

Cause 4: the people doing the work were told, not asked

The support team found out about the AI programme when the vendor showed up. From that moment, every error became evidence for a case they were already building, and nobody flagged the near-misses that would have made it better.

This is not sabotage, it is entirely rational self-preservation, and it is the failure mode that technical teams underestimate most. The people who know which fifteen actions make up 90% of your volume are the ones being automated. You need them in week one, deciding what to automate first, and you need an honest answer to "what happens to my job" before they have to ask.

The honest answer is usually good: nobody's job is 4,000 hours of pasting tracking numbers, and removing that part is not removing the person. Say it out loud, early, and mean it.

The check: who from the affected team is in the room on day one, and what were they told?

Cause 5: one bad incident, no containment story

The system did something wrong and visible. There was no recall mechanism, no ledger to reconstruct what happened, no per-action-type kill switch — so the only available lever was switching the whole thing off. Nobody ever switched it back on.

Every deployment has an incident. The ones that survive have three things ready before it happens: an undo, a log that explains, and a way to disable exactly one action type without touching the rest.

The check: walk me through the incident. It executed the wrong thing on 40 customers. What do I click, what do I know, and what stays running?

The pattern underneath all five

Every one of these is an operations failure wearing a technology costume. Not one of them is about model quality. That is why buying a better model does not fix a dead pilot, and why the second attempt usually dies the same way as the first.

The five questions, in one place 1. What action does it take, unaided, in week four?
2. What baseline are we taking before we start?
3. What is the numeric success condition for automation one?
4. Who from the affected team is in the room, and what were they promised?
5. Show me the incident: undo, log, and per-action kill switch.

A vendor with good answers to all five may still fail you. A vendor with bad answers to any of them will.

Keep reading

This is what we do all day.

Support automation, AI phone agents, n8n back-office work, and the engineering loop itself — always behind a gate you control.