Guide 01

AI customer support, honestly described.

Most of what gets written about AI customer support is written by people selling it, which makes it hard to work out what actually happens when you put one on a website. This is the version we would give a customer on a call: what a bot genuinely takes off the queue, what it should never touch, and how to tell within a month whether it is earning its place.

The only number that predicts whether this works

Before evaluating any tool, go and read your last two hundred support conversations and sort them into two piles.

The first pile is lookups. What are your hours. Do you take my insurance. Where is my order. How much is a deep clean. Can I book Thursday. Do you serve my zip code. These have one correct answer that does not change based on who is asking, and the person answering them is reading it off a page or a screen.

The second pile is judgment. My order arrived broken and I am angry. Can you make an exception. Which of these two options is right for my situation. I want to cancel and I want you to talk me out of it. These require someone to weigh something, and frequently to take responsibility for a decision.

The ratio between those piles is the honest ceiling on what automation can do for you. A restaurant, a clinic, a home services company and an ecommerce shop typically find the first pile is most of the volume, because the questions arriving at a small business website are overwhelmingly logistical. A consultancy or a complex B2B product often finds the opposite. Neither is a better business; they just have different inboxes.

Vendors quote deflection rates as though they were a property of the software. They are mostly a property of your inbox. Anyone quoting a number before they have seen your tickets is guessing, including us.

If you want the arithmetic rather than the principle, we have gone through it twice at length: once on how much of a ticket queue a chatbot actually removes, and once on which categories of ticket automate cleanly and which do not.

What a support bot is genuinely good at

Sorted roughly by how reliably it works, rather than by how impressive it sounds in a demo.

Two of these deserve their own treatment. Automating the questions you answer every day is the fastest of them to get right, and covering the hours nobody is staffed for is usually the one that changes a customer's opinion of you, because the alternative they are comparing it against is silence.

What it should not be doing

This list matters more than the one above, because the failures are more expensive than the wins.

Handoff is the part that decides whether people trust it

Everything else on this page is secondary to getting handoff right. A bot that answers eight questions well and traps you on the ninth is remembered for the ninth.

Four triggers that should always produce a human, without negotiation:

  1. The customer asks. "Talk to a person", "agent", "is anyone there". Immediately, first time, no "let me try to help first". Asking a second time is where goodwill dies.
  2. Frustration appears. Repetition, capitals, swearing, or the same question rephrased twice. People rephrase when they think they are being misunderstood, and they are usually right.
  3. The topic is on the do-not-automate list. Refunds, complaints, cancellations, anything regulated.
  4. Two failed attempts. If the bot has missed twice, a third attempt will not land, and the customer has already concluded it cannot help.

How the handoff feels matters as much as when it fires. The customer should not have to repeat themselves; whoever picks up should arrive with the transcript. If nobody is available, say so plainly, give a real timeframe, and take a way to reach them back. "Our team is offline until 9am, leave your number and you will get a call before ten" is a good outcome. Silence after a promise is worse than never having offered.

Where the answers come from

A support bot is only as good as what it has been told, and this is where most of the ongoing work lives.

The starting set is smaller than people expect. Your top ten questions, answered the way your best staff member answers them, covers a surprising share of the volume. Do not begin by writing two hundred answers; begin by writing ten and watching what the bot gets asked that it cannot handle.

Then it becomes a reading job. Once a week, read the conversations where the bot failed or handed off, and ask a narrow question: was this a gap in what it knows, or a question it should never have been answering? Gaps get filled. The second category gets added to the handoff rules. Programs that do this weekly for a month end up somewhere good. Programs that deploy and walk away stay exactly as useful as they were on day one, which is usually "occasionally".

One caution about tone: an answer that is technically correct and reads like a policy document will get a follow-up question, which means it did not actually deflect anything. Write the answers the way a helpful colleague would say them out loud.

Businesses running an older ticketing setup often find this step is really a question about the tool underneath it, which we have written about separately in modernizing a help desk without replacing everything at once.

Measuring it without fooling yourself

Deflection rate is the number every dashboard leads with and the easiest one to make look good. A bot that is hard to escape deflects beautifully and is quietly costing you customers. Read these four together instead.

Alongside the numbers, read twenty transcripts a week yourself. Metrics tell you the shape of the problem; transcripts tell you what is actually going wrong, and they routinely surface things no dashboard would have flagged.

What you feed it and how you keep it current is its own subject, and a bigger one than it looks: the guide to chatbot knowledge and measurement covers curating sources, where wrong answers actually come from, and the numbers worth reading afterwards.

A four-week start that does not overreach

  1. Week one: read the inbox. Sort two hundred conversations into lookups and judgment. Write down your top ten questions and the answer to each in your own voice. Decide the do-not-automate list now, while nothing is live.
  2. Week two: ship narrow. The ten answers, a clear route to a human, and honest out-of-hours language. Put it on the pages where those questions actually arrive rather than site-wide.
  3. Week three: read what broke. Every failed and handed-off conversation. Fill the gaps, tighten the handoff triggers, fix the answers that produced a follow-up question.
  4. Week four: widen carefully. Add the next tier of questions and only then consider more pages. Start reporting the four numbers above so there is a baseline before anyone asks whether it is working.

The failure mode this ordering avoids is the common one: launching with an enormous knowledge base, no handoff discipline, and no plan to read the transcripts. That version answers a lot of questions slightly wrong and teaches customers to skip the widget.

Three failure patterns worth recognizing early

These account for most of the support bots that get quietly switched off six months after launch.

The bot that will not let go. Someone built it to maximize deflection, so every route to a human is buried behind "let me try to help with that first". Deflection metrics look excellent. Customers learn that the widget is a wall and start emailing the sales address instead, which is now your support queue and nobody is measuring it. The tell is a falling chat volume alongside rising contacts through channels you did not intend.

The bot nobody reads. It launched, it works about as well as it did on day one, and no human has opened a transcript since. It answers the questions it was given and fails silently at everything else. This one is not harmful, it is just an asset depreciating in public. The fix is a standing thirty minutes a week, which is genuinely all it takes.

The bot that answers too confidently. Asked something outside its knowledge, it produces a plausible answer rather than admitting the gap. In support this is worse than silence, because the customer acts on it and the correction arrives later and angrier. A bot that says "I do not have that, let me get someone who does" is doing its job. Test for this deliberately before launch by asking it things you know it was never told.

What it costs, and what it saves

The software cost is the easy half and usually the smaller one. The real inputs are the afternoon of setup, the thirty minutes a week of reading, and the honest fact that someone has to own it. A bot with no owner drifts into the second failure pattern within a quarter.

On the saving side, be careful which number you claim. Hours saved is the one people reach for and it is the softest, because the time freed rarely converts neatly into anything measurable. Two firmer numbers:

If your inbox is mostly judgment calls, neither number will be large, and the honest conclusion is that this is a small improvement rather than a transformation. That is a perfectly reasonable place to land, and it is better to know in week one than after a rollout.

Where support sits next to everything else

Support automation is usually the first thing a business puts a bot on, because the pain is obvious and the questions repeat. It is worth knowing that the same widget does two other jobs, and that the boundaries between them are blurrier than the categories suggest.

The conversation that starts "do you take my insurance" is a support question and also, quite often, a new patient deciding where to book. The one that starts "where is my order" is support, and the reply is a chance to keep a customer who was one bad experience from leaving. Treating the widget as purely deflection means measuring only half of what it does, and it is the half that does not show up in revenue.

Two practical consequences. First, do not put the support bot only on a help page. The questions arrive on service pages, pricing pages and product pages, because that is where people are when they hesitate. Second, when a support conversation turns into intent to buy or book, the bot should be able to carry it rather than answering the question and stopping. Capturing that is a different job with its own rules, and it is covered separately in our guide to chatbot lead capture.

The reverse also holds: a lead-capture bot that cannot answer a basic logistical question loses the lead at exactly the moment it mattered. In practice most small businesses want one conversation that can do both, which is less complicated than it sounds because both jobs draw on the same ten answers you wrote in week one.

Two adjacent questions come up often enough to have their own answers. On whether to run automation or staffed chat, see chatbot or live chat, and when each is the wrong choice. On the commercial side of answering faster, see what support response time does to churn.

Common questions

What percentage of tickets can AI actually deflect?

It depends on what your tickets are, not on how good the bot is. An inbox dominated by hours, pricing, order status and booking questions deflects a lot, because those are lookups. An inbox of judgment calls deflects very little and should not try. Sort your last two hundred tickets before you estimate anything.

Will customers be annoyed by a chatbot?

They are annoyed by one that will not let them reach a person, which is a different thing. A bot that answers instantly beats a form, and one that hands off in a single step beats a phone queue. The complaint is about being trapped, not about software.

When should a chatbot hand off to a human?

On request, immediately. On any sign of frustration. On money, cancellations, complaints, and anything regulated. And after two failed attempts, because the third will not go better.

Does AI customer support replace support staff?

It replaces the repetitive part of the queue, which is the part staff least want. Treating it as a headcount cut tends to make support worse, because what remains is the hard conversations and nobody has slack for them.

How long before it is useful?

A bot covering your top ten questions is useful the day it ships. Getting it genuinely good takes a few weeks of reading transcripts. The long tail is never finished, and past a point the remaining questions are ones a person should answer anyway.

Start narrow

Ten answers and a clear route to a human.

That is a working support bot, and it is an afternoon of setup. Widen it once you have read a week of transcripts.