← Playbook

“We'll just use ChatGPT.” A post-mortem.

It works brilliantly for one person. It falls apart the moment it becomes a process. Here is exactly where the seam is, and why it is not an argument against using it.

2026-07-31/6 min read/Levelbrook AI Practice

Somebody in your company has already automated part of their job with a chat window. They paste in the customer email, get a reply, tweak it, send it. It genuinely saves them time and they are right to do it.

Then somebody says: let's roll this out to the whole team. And that is where it comes apart — not because the model is bad, but because a person using a tool and an organisation running a process are different things with different requirements.

What breaks, specifically

It does not know your policy, so it invents one. Asked about a refund window it has never been told, it will produce a confident, plausible, industry-standard answer. Thirty days is a very common refund window. It may not be yours. The individual using it catches this because they know the policy. The process does not catch it, because the whole point of scaling was that the next person does not have to know.

Everyone's version is different. Six people build six private prompts and get six voices, six policy interpretations, six standards for when to escalate. You have not standardised anything — you have added variance and made it invisible.

There is no record. A customer asks in November why they were told something in August. The answer lived in a chat window that belonged to an employee who has since left. There is no ledger, no evidence, no way to reconstruct the decision.

Nothing improves. When it gets something wrong, the person fixes it in the moment and moves on. That correction is the most valuable data your company generates and it evaporates every single time.

It still cannot do anything. The reply gets written in one window and typed into another. The refund is still issued by hand. The address is still changed by hand. You have automated the sentence and left the work.

And the data question is real. Pasting customer records into a general consumer tool is a different compliance posture than an enterprise deployment, and "we didn't know people were doing that" is a bad sentence to say to an auditor.

The seam, stated precisely

The distinction that matters A chat window is an interface. What you need is a system. The model can be identical. The difference is grounding, action, record and feedback — and none of those live in a chat window.

A system has four properties an individual's chat window structurally cannot:

  1. Grounding — it answers from your policies, your tickets, your product data, with citations, and refuses when it cannot find them.
  2. Action — it produces an executable, validated, reversible proposal, not a paragraph somebody retypes.
  3. Record — every proposal, approval, rejection and execution is logged with its evidence.
  4. Feedback — corrections are captured with a reason and used, rather than being lost in somebody's browser history.

Why this is not a criticism of the person who did it

Genuinely: the employee who wired up their own workflow is the most valuable person in this story. They have done free, high-quality discovery. They know which tasks are worth automating, which prompts work, where the model reliably fails, and which of their colleagues' objections are real.

Every engagement we run starts by finding those people. They are usually not in management, they are usually slightly sheepish about it, and their private prompt library is the best specification document in the building.

What to do on Monday

  1. Find out who is already doing it. Ask without any threat attached — you want the honest answer, not compliance.
  2. Collect the prompts. That is your requirements document, written by the people who do the work.
  3. Write the data rule now. Not a ban — bans just push it underground. A clear statement of what may be pasted where, and an approved tool that is actually good enough to use.
  4. Pick the one task with the highest volume and lowest blast radius, and build it properly. Grounded, with an action, behind an approval queue, with a log.

"We'll just use ChatGPT" is not the wrong instinct. It is the right instinct, stopped one step too early — at the sentence, instead of at the work.

Keep reading

This is what we do all day.

Support automation, AI phone agents, n8n back-office work, and the engineering loop itself — always behind a gate you control.