Guide 04

What it knows, what you read back, and what you owe the customer.

A chatbot is three things running together: a body of knowledge, a record of what happened, and a promise about how that record is handled. Get the first wrong and it answers confidently and incorrectly. Get the second wrong and you cannot tell. Get the third wrong and it stops being a technical problem.

Knowledge is curation, not upload

The instinct when setting up a chatbot is to give it everything: the site, the PDFs, the old FAQ, the internal wiki, the price list from the shared drive. It feels thorough. It is the single most reliable way to end up with a bot that tells customers things that used to be true.

Understand what actually happens. The bot does not know which of your documents is authoritative. When two sources disagree, it does not arbitrate; it answers from whichever looks most relevant to the question asked. A superseded policy sitting in the pile is not inert. It is a confident wrong answer waiting for the right question.

So the discipline is subtraction. A tight set of current material beats a large set that is partly stale, every time, and it is not close. Three practical rules:

What belongs in there instead is the unglamorous material described in the build guide: the replies you already send repeatedly, the specifics your marketing copy avoids, and the awkward questions with an answer you have decided on in advance rather than one the bot improvises.

The mechanics underneath this are worth understanding before you decide what to feed it: how a chatbot knowledge base actually retrieves an answer explains why relevance, not recency, decides which document wins.

Wrong answers, and where they actually come from

When a chatbot says something false, the assumption is that the model made it up. Occasionally true. Far more often the bot faithfully reported something you gave it, and the thing you gave it was old.

The distinction matters because the fixes are different. Fabrication is handled by constraining what the bot can draw from and by giving it explicit permission to say it does not know. A bot with no route to admit ignorance will always produce something, because producing something is the only behavior available to it. "I am not sure, let me get you to somebody who is" has to be a first-class outcome, not a fallback nobody configured.

Staleness is handled operationally, and there are only two habits that work. When a price, policy, service or opening hour changes, the source material changes in the same sitting. And once a quarter, somebody asks the bot the five questions that matter most commercially and reads the answers properly. Fifteen minutes, and it catches nearly everything.

One more category worth naming, because it is the one people miss: the confidently unhelpful answer. Not wrong, exactly, just useless. "Our pricing depends on your requirements." Technically accurate, and the visitor leaves. These do not show up as errors anywhere, and the only way to find them is to read transcripts.

Languages, if you need them

Multilingual support is one of the few genuinely easy wins here, and it is worth being clear about where the difficulty actually sits. Answering in Spanish or Polish is close to free. The hard part is everything downstream: if a Spanish-speaking customer has a conversation the bot cannot finish, and the person who picks up the escalation only speaks English, you have built a better front door onto the same wall.

So decide the handoff before you switch languages on. If you can serve those customers when the conversation gets hard, enable it. If you cannot, it is more honest to say which languages you support than to have a bot that converses warmly right up until a person is needed.

On the language question specifically, running a chatbot in more than one language covers what changes operationally rather than technically.

The numbers worth reading

Chatbot dashboards are generous with metrics, and most of what they show best is what matters least. Conversations started is the clearest example: it goes up when you make the widget more aggressive, which is a setting, not an achievement.

Four numbers carry almost all of the signal.

  1. Resolution rate. The share of conversations that ended with the person's question answered and no human involved. This is the number that says whether the knowledge is any good. It is also the one most often calculated dishonestly: a conversation where somebody gave up and closed the window is not a resolution, and any tool that counts it as one is not worth reading.
  2. Escalation rate, with reasons. The raw percentage matters less than the breakdown. Escalations clustered on one topic mean a specific gap you can close this afternoon. Escalations spread evenly mean the bot is scoped too broadly and is attempting work it should be handing over immediately.
  3. Outcomes. Leads captured, appointments booked, orders tracked. Whatever you decided the one job was. This is the only figure that settles an argument about whether the channel earns its place.
  4. Response reality on escalations. Not a bot metric at all, which is why nobody watches it. What share of handed-off conversations got a human reply, and how fast. A high escalation rate with fast replies is a working system. A low escalation rate with replies nobody sends is a machine for losing customers politely.

Satisfaction scores deserve a caveat. Collect them, but weight them lightly: the people who rate a chat interaction are disproportionately the ones with strong feelings, and the silent majority who got what they wanted and left are invisible in that number. A thumbs-down with a comment attached is worth more than fifty ratings without one.

And the number nobody reports on: how many people opened the widget, read the opener, and closed it without typing. That is a message about your opener, and it is usually a large number.

There is more to say about reading the reporting itself, which what chatbot analytics tell you about your customers takes further than the four numbers above.

Transcripts beat dashboards

Everything above tells you where people struggled. Only the transcripts tell you why, and the why is almost always more mundane than expected. A question phrased in a way that reads as intrusive. An answer that assumed knowledge the visitor did not have. A word your industry uses that your customers do not.

The routine that works is short enough that it actually gets done. Twenty conversations a month, chosen at random rather than by rating, because the selected ones are either the disasters or the successes and the interesting material is in the middle. List every answer that was poor. Fix the top three. Stop.

Reading at random matters more than it sounds. Sorting by satisfaction shows you the tails. The bulk of your traffic sits in the unremarkable middle, where somebody got a technically adequate answer and quietly decided you were not worth the trouble.

If the numbers are telling you the bot is attempting work it should be handing over, that boundary is drawn in the AI customer support guide, which covers what a bot should never touch and how the handoff ought to read.

Working out whether it paid for itself

Chatbot ROI cases tend to be built out of vendor arithmetic and should be treated with suspicion, including when the vendor is us. Two numbers can be calculated honestly.

Time saved. Resolved conversations, multiplied by the minutes each would have taken a person, multiplied by a loaded hourly rate. Be conservative about the minutes: a question answered in chat would often have been answered in a two-line email, not a ten-minute call. This number is usually smaller than the marketing suggests and more defensible than anything else on the page.

Revenue influenced. Leads captured through chat, multiplied by your close rate, multiplied by average value. Then discount it, heavily, because a share of those people would have contacted you through the form anyway. Attribution here is genuinely hard and anyone claiming precision is selling something. Halving it is a defensible instinct.

There is a third effect that is real and resists measurement: the customers who got an answer at 10pm and did not go and look at a competitor. You cannot count them. It is fine to believe the effect exists and to keep it out of the spreadsheet.

We have worked a fuller version of this calculation, including the assumptions worth arguing with, in measuring the return on AI customer service.

The same knowledge, used twice

Once you have done the work of writing down real answers to real questions, you own something more useful than a chatbot. You own the content of a help center, and most businesses never notice.

The overlap is close to total. A well-curated chatbot knowledge base is a set of specific questions with plain answers, which is exactly what a self-service portal is, and exactly what search engines index well. The same forty answers can sit behind the widget, live as browsable pages, and appear in results when somebody searches the question at midnight without ever visiting your site.

That third use is the one worth taking seriously, because it changes what the writing is for. An answer written only for a chat window can be terse and assume the visitor is already on your pricing page. An answer that also has to work as a page needs enough context to stand alone. Writing to the second standard costs a little more and gives you both.

Two cautions. Publishing every internal answer verbatim will expose material you meant for a one-to-one conversation, so decide per answer rather than in bulk. And a published help page inherits the staleness problem with an audience: a wrong chat answer reaches one person, a wrong page reaches everybody who searches for it, indefinitely.

Taken far enough, this becomes a help center rather than a chat feature, which is the subject of building a self-service portal customers actually use.

Who owns this, in practice

Every recommendation on this page assumes somebody is doing it, and the most common reason none of it happens is that the chatbot belongs to whoever set it up, which after a few months is frequently nobody.

The ownership needed is smaller than it sounds. One named person, thirty minutes a month, with three standing responsibilities: read twenty random transcripts, update the knowledge whenever something about the business changes, and confirm escalations are still reaching a human. That is the entire job, and it does not require the person who built the thing.

What it does require is that the person hears about changes. The failure is almost never neglect, it is that pricing changed in a meeting the chatbot owner was not in. Whatever process you use to announce a price change, a new service or a change in hours, the knowledge base goes on that list, next to the website and the phone greeting.

Trust, which is the part with consequences

A chat window feels casual to the person typing into it, and that is exactly why it collects more than people intend. Visitors put things in chat that they would never put in a form: medical detail, financial circumstances, complaints about a named employee, occasionally a password. Your obligations follow the sensitivity of what is actually said, not what you designed the field for.

A short list that covers most of it.

None of this is a reason not to run a chatbot. It is a reason to decide these things deliberately at setup, when they take an hour, rather than during an incident, when they take considerably longer and somebody else sets the terms.

There is one more thing to settle in advance, and it is the only item here with a deadline attached: what you do when the bot gets something wrong in front of a customer. It will happen, and the difference between a small problem and a lasting one is entirely in the first hour.

Decide now who can turn it off, and make sure that is more than one person and does not require the vendor. Decide what you say, which is some version of the plain truth: the answer was wrong, here is the correct one, here is what we are doing about it. And decide what you honor. If a bot quoted a price it should not have, paying that price once is usually cheaper than the argument, and considerably cheaper than the review.

Then fix the source rather than the symptom. The instinct is to add a rule preventing that exact answer. The better move is to find the stale document that produced it, because there is rarely only one wrong fact in a file that has gone unmaintained long enough to cause an incident.

The security and privacy side has its own treatment in what a business is responsible for when a chatbot handles customer data, which goes further into the regulated cases than is sensible here.

Common questions

What should go in a chatbot knowledge base?

The replies you already type repeatedly, plus the specifics your site avoids: real ranges, turnaround times, service areas, what you do not do. Curate rather than dump. Stale knowledge does not fail silently, it fails fluently.

How do you stop wrong answers?

Constrain the sources, give the bot explicit permission to say it does not know, and remove outdated material instead of leaving it in place. Then read transcripts. Most wrong answers are old facts, not invention.

Which metrics actually matter?

Resolution rate, escalation rate with reasons, the outcome you built it for, and how fast humans reply to escalations. Conversations started rises when you make the widget pushier, which is a setting rather than a result.

How do you calculate ROI?

Time saved, being conservative about minutes, plus revenue influenced, discounted heavily because some of those people would have contacted you anyway. Anyone claiming precise attribution here is selling something.

What should you tell customers about their data?

That it is a bot, what is stored, for how long, and who sees it, in the opener rather than in a policy. And do not collect sensitive categories in chat unless you have the handling to match.

Knowledge and reporting

See what Alma stores, surfaces, and hands back.

Knowledge sources, escalation rules and the reporting that tells you whether any of it is working.