When a chatbot is the wrong answer
Half the businesses that ask us for a chatbot have a documentation problem, not a conversation problem.
The short version
- A bot can only answer from what exists. Undocumented policy becomes invented policy.
- Deflection is not resolution — the gap between the two is where churn hides.
- Write and publish your twenty most-asked answers first. Often the bot stops being urgent.
- Where bots do work, it is because they are wired to a real system, not to a vibe.
A chatbot can only answer from what exists. If your policies live in three people’s heads, a bot will confidently invent the rest, and you will have created a new problem with a friendly interface on it.
Half the businesses that ask us for a chatbot have a documentation problem, not a conversation problem. That is not a criticism — it is the normal state of a company that grew by knowing things rather than writing them down.
Deflection is not resolution
The metric most bots are sold on is deflection: the share of enquiries that never reached a human. It is a flattering number, because an abandoned conversation and a confidently wrong answer both count as successes.
Gartner's research is blunt about the gap. AI deflects more than 45% of customer queries, but only around 14% of issues reach genuine self-service resolution — the other thirty-odd points are customers who came back through a different channel, or gave up. Even for issues customers themselves described as very simple, the resolution rate reached only 36%.
A bot with a 90% deflection rate can sit on a 40% resolution rate. The number worth running the business on is the one that confirms the problem went away.
When self-service fails, the causes are consistent and they are not model quality: around 43% of failures happen because the customer cannot find content relevant to their issue, and about 45% because the system did not understand what they were trying to do. Both are content and scope problems.
What we suggest instead
Write the twenty questions your customers actually ask, answer them properly once, and put them somewhere a person can find. Concretely:
- Pull the last two hundred enquiries from email, WhatsApp and the phone log. Do not guess the list — the real one is always different from the imagined one.
- Cluster them. You will usually find that twelve to twenty question types cover eighty percent of volume.
- Write one clear answer each, approved by whoever actually owns that policy. Include the awkward ones: refunds, delays, pricing exceptions.
- Publish them where customers already look, and where staff can copy from them.
- Measure again in a month. Enquiry volume usually falls enough that the bot is no longer the urgent purchase.
If it still is, now it has something true to work from — and the deflection numbers move accordingly. Teams whose help content was updated in the last thirty days consistently outperform teams whose content has not been reviewed in six months, by more than a factor of two.
Where bots do earn their keep
- After-hours triage. Not answering everything — capturing the enquiry properly, setting an expectation, and routing urgent cases to a human path that genuinely exists.
- Status lookups against a real system. "Where is my order" is a database query wearing a conversation. Wire the bot to the order system through an API and it will be right every time; let it guess and it will be wrong at the worst moment.
- Internal assistants over your own documents. The highest-value deployments we have built are for staff, not customers: searching procedures, contracts and past quotes that nobody can find by filename. The tolerance for an imperfect answer is higher, and the person asking can tell when it is wrong.
If you do build one, build it properly
- Ground every answer in retrieved source material, and show the source. If the retrieval finds nothing, the correct behaviour is to say so and hand over — not to improvise.
- Define the scope narrowly and refuse politely outside it. Bots that try to answer everything score worse than bots that answer six things reliably.
- Make escalation one click, visible, and always available. Hiding the human is how you convert a support cost into a lost customer.
- Log every conversation and read them weekly for the first two months. The transcripts are the best product research you will ever get.
- Decide the language question up front. In Mauritius that usually means English and French at minimum, and being honest about how the bot handles Kreol — including recognising it and switching to a person if the answer quality is not there.
Measure it honestly
Two numbers, reviewed monthly. Containment: the share of conversations that ended without escalation. True deflection: containment minus anyone who came back through another channel within seven days. The second number is typically thirty to forty percent lower than the first, and it is the only one worth reporting to a board.
Set a floor as well as a target: if containment is below thirty percent after a month, the answer is not a better model. It is missing content.
Frequently asked questions
Will a chatbot reduce our headcount?
In a business under fifty people, almost never. What it changes is what the same people spend their day on — the same pattern we described in the invoice automation write-up.
Can it use our existing documents?
Yes, and that is the right approach — retrieval over your own material rather than a model answering from general knowledge. The quality ceiling is set by the documents, which is why we start there.
What does it cost to run?
Per-conversation costs are small; the ongoing cost that matters is the half-day a month someone spends reading transcripts and updating answers. Budget for that or the bot will decay.
Can you just tell us whether we need one?
That is what the free review is for. An hour, an honest answer, and you keep the write-up either way. Book a call.
Sources
- Gartner customer service research, 2025 — self-service resolution rate (14%), deflection above 45%, and causes of self-service failure.
- Gartner 2025 Customer Service Technology Survey — containment benchmarks for retrieval-based versus rule-based deployments.
- HubSpot State of Service and industry deflection benchmark syntheses, 2025–2026 — effect of knowledge base freshness on deflection.
