It works brilliantly for one person. It falls apart the moment it becomes a process. Here is exactly where the seam is, and why it is not an argument against using it.
Somebody in your company has already automated part of their job with a chat window. They paste in the customer email, get a reply, tweak it, send it. It genuinely saves them time and they are right to do it.
Then somebody says: let's roll this out to the whole team. And that is where it comes apart — not because the model is bad, but because a person using a tool and an organisation running a process are different things with different requirements.
It does not know your policy, so it invents one. Asked about a refund window it has never been told, it will produce a confident, plausible, industry-standard answer. Thirty days is a very common refund window. It may not be yours. The individual using it catches this because they know the policy. The process does not catch it, because the whole point of scaling was that the next person does not have to know.
Everyone's version is different. Six people build six private prompts and get six voices, six policy interpretations, six standards for when to escalate. You have not standardised anything — you have added variance and made it invisible.
There is no record. A customer asks in November why they were told something in August. The answer lived in a chat window that belonged to an employee who has since left. There is no ledger, no evidence, no way to reconstruct the decision.
Nothing improves. When it gets something wrong, the person fixes it in the moment and moves on. That correction is the most valuable data your company generates and it evaporates every single time.
It still cannot do anything. The reply gets written in one window and typed into another. The refund is still issued by hand. The address is still changed by hand. You have automated the sentence and left the work.
And the data question is real. Pasting customer records into a general consumer tool is a different compliance posture than an enterprise deployment, and "we didn't know people were doing that" is a bad sentence to say to an auditor.
A system has four properties an individual's chat window structurally cannot:
Genuinely: the employee who wired up their own workflow is the most valuable person in this story. They have done free, high-quality discovery. They know which tasks are worth automating, which prompts work, where the model reliably fails, and which of their colleagues' objections are real.
Every engagement we run starts by finding those people. They are usually not in management, they are usually slightly sheepish about it, and their private prompt library is the best specification document in the building.
"We'll just use ChatGPT" is not the wrong instinct. It is the right instinct, stopped one step too early — at the sentence, instead of at the work.
Support automation, AI phone agents, n8n back-office work, and the engineering loop itself — always behind a gate you control.