Customer service automation that doesn't turn your company into a call center

When using this phrase, Google asks directly whether automation will improve service quality. The first ten results provide a list of examples, but none of them answer the question. We show which queries can be safely automated, where to set the threshold for handing them over to a human, and what happens in the first few weeks after launch.

·5 min read·Jakub Dulas
Customer service automation that doesn't turn your company into a call center

The biggest concern with automating customer service isn’t about money. It goes like this: a customer will send a message, get some gibberish in response, and go to the competition. This concern is valid, because that’s exactly what happens when you tell the system to respond to everything. The trick is to have it respond to only a few things.

Below is a breakdown of inquiries into those that can be handled without risk and those that should never be handled by the system. Plus, the threshold at which the conversation is handed back to a human.

Which queries can be handled without risk

There’s only one criterion. The answer must always be the same and must be verifiable in your system, without having to ask anyone for their opinion.

This group includes order status, business hours, return policies, product availability, how to file a complaint, and bank transfer details. Anything where the answer is based on a fact stored in your system.

In a typical service company, this accounts for 60–70 percent of the inbox volume. Not 100 percent. And that’s a good thing, because the remaining thirty percent is left for people who finally have time to handle them.

Never Automate These

Three categories where an automated response does more harm than good.

Complaints and grievances. An upset customer wants to speak to a real person, not a polished script. An automated system can accept the report and assign a case number, but it shouldn’t resolve the issue.

Unusual or first-of-its-kind cases. If a question is asked for the first time, the system has nothing to base its response on and will start guessing.

Anything involving the customer’s money in a non-obvious way. Invoice adjustments, individual discounts, disputed payments. Here, the cost of a single wrong answer outweighs the savings from a hundred correct ones.

Where to set the threshold for transferring a query to a human

This single setting determines the success of the entire implementation, and yet it almost never appears in the product offering or in the demo.

A good system has a confidence threshold for responses. Below the set threshold, it doesn’t respond but instead forwards the matter along with the entire conversation history. Start with a high—that is, conservative—threshold. It’s better for the system to forward too many conversations at the start than to give a single confident but incorrect response.

The second threshold concerns the customer. If someone rephrases the same question twice, it means they didn’t get an answer. Forward the conversation without asking for their consent.

The third trigger is emotions. Detecting agitation in the content should immediately end the conversation with the bot.

What the system does when it doesn’t know the answer

Ask every provider about this before you sign a contract.

There are three possible behaviors. The system says outright that it doesn’t know and forwards the request. The system gives an evasive answer and pretends to know. Or the system makes up a specific detail that isn’t in the data.

Only the first option is acceptable. The third scenario actually happens if the agent doesn’t have strict access restrictions to the company’s knowledge base, and then the customer receives a specific number or deadline that no one in your company has ever set. Ask to see this behavior demonstrated during a demo, using a question outside the scope. The vendor’s response will tell you more than the entire presentation.

Should you tell the customer they’re chatting with a machine?

Yes, and not just because the AI Act requires it for systems that interact with people.

A client who is informed right away asks questions more simply and gets to the point faster. A client who figures it out on their own halfway through the conversation feels deceived, and that’s the moment when the company loses more than it would from ten delayed responses.

All it takes is one sentence at the start, plus a clear option to transfer to a human agent.

How long does it take for quality to return to pre-implementation levels?

The honest answer is: things will be worse for the first few weeks.

The first two to four weeks are for calibration. The system gets it right 70 percent of the time; someone reviews the conversations and makes corrections. This is normal and needs to be planned for, not discovered later.

The pre-implementation level usually returns after a month. Quality exceeds pre-implementation levels in the second or third month, mainly thanks to response time, since the automated system responds in seconds—even at night.

If a provider promises full quality from day one, ask who will be reviewing the conversations during the first month. Someone has to do it.

What happens when you contact us

We start by reviewing the last 100 inquiries from your inbox and calculating what percentage is suitable for handoff. This is a specific number, not a claim.

The owner speaks with you; the call lasts 45 minutes. If the numbers show that fewer than half of the calls are repetitive, we’ll let you know and advise against implementation at this stage.

The first phase covers one channel, not all at once. The team’s subscription starts at 3,875 PLN per month, and you can calculate the cost using the calculator before the call. The guarantee covers repairs at no extra charge.

Questions & Answers(FAQ)

By reducing response times and freeing people from having to answer repetitive questions. The chatbot responds within seconds—even outside of business hours—to inquiries about status, terms, and availability. This gives the team time to focus on complex issues that previously had to wait in line behind questions about business hours.

Partially. In a typical service company, 60–70 percent of inquiries always have the same answer based on the data in the system, and those can be delegated. Do not delegate complaints, unusual cases, or disputed settlements.

He should say outright that he doesn’t know and refer the matter to someone else, along with a summary of the conversation. The other two responses—a evasive answer and a made-up detail—render this solution unacceptable. Ask the vendor to demonstrate this in a demo using a question outside the scope.

Yes. The AI Act requires that systems interacting with people be labeled as such, and even aside from that requirement, it simply pays off. An informed customer asks simpler questions. A customer who figures it out on their own halfway through the conversation feels deceived.

The first two to four weeks are for calibration, and during this time, quality may be lower than before implementation. The baseline level usually returns after a month, and performance begins to improve in the second or third month. Someone needs to review the conversations during the first month.

Got a similar process on your side?

If something in this article sounds like your day-to-day - let's talk. We'll tell you plainly what can be improved, and what's not worth touching.