Written for the person who has to sign off on it, not for the person who has to build it. No vendor pitch, no benchmark charts, and the failure numbers included.
Not the ones in the headlines. The real order is set by volume and reversibility, and it is remarkably consistent across every industry we walk into.
Everybody publishes the accuracy number. Almost nobody publishes the failures. The failures are the entire product, because they tell you exactly which gate to build.
Everyone assumes a review gate is a tax on throughput. Measured against the task it replaces, it is the cheapest thing in the building — but only if you build the interface properly.
Not answering hard questions. Not calming angry customers. Typing. Here is how to get the actual number for your own company in about three days, before you spend a dollar on AI.
A surprising share of what gets logged as a model failure is the system faithfully reporting a conflict that was already in the company. Grounding does not just reduce errors — it audits you.
Service businesses obsess over lead cost and ignore the leads they already paid for and then dropped. The arithmetic is brutal and almost nobody has run it.
Any team that has actually run this in production has a list of what the system got wrong, sorted by category, with counts. Asking for it takes ten seconds and separates the operators from the demos.
Accuracy is the wrong headline number. The number that decides whether a system is deployable is how well it knows what it does not know.
Different industries, different words, same fifteen jobs. If you recognise more than eight of these, you have a straightforward automation project and you do not need anyone to tell you that.
The model is a commodity that gets better every quarter without you. The reviewable action pipeline is the thing that is actually hard, actually yours, and actually worth money.
The model is a commodity that improves on somebody else's budget. The thing your competitors cannot copy is the twenty years of operational data sitting in the system you complain about.
They do not fail for technical reasons. They fail in five specific, predictable, entirely preventable ways, and you can check for every one of them before you sign anything.
Support organisations feel infinitely varied from the inside. From the outside they are a very short list, repeated, with different names attached.
It works brilliantly for one person. It falls apart the moment it becomes a process. Here is exactly where the seam is, and why it is not an argument against using it.
The thing that gets an operations director, a compliance officer and a nervous CFO to say yes is not the model. It is being able to answer, in eleven months, what happened and why.
None of them are technical. All of them are answerable by anyone who has actually deployed this in a real company. Print them out and bring them to the demo.
We build a great deal in n8n and recommend it constantly. It is also the wrong tool for about a third of what people put in it, and the failure is always the same one.
Not a prompt. Eight structural controls, none of which depend on the model behaving well, and all of which you can build before writing a single line of agent code.
If after reading it you want somebody to build the thing, that is what we do.