Blog

When customers can’t tell it’s AI, quality assurance can’t be an afterthought 

Daniel Gil, AI CX Specialist
3rd September 2026

Much of the regulation now shaping how we deploy AI in customer service — the multi-language requirements, the mandatory route to a human on request — rests on a quiet assumption: that today’s AI still behaves like the IVRs and bots of the last fifteen or twenty years.  

In most companies, that assumption is still fair. But it’s ageing, and fast, and it raises an awkward question about what we measure and why. 

For a while now, the things that gave a virtual agent away were obvious. The voice. The tell-tale latency of the technology. And, above all, the way we made it work — narrow, scripted, unmistakably a machine. We worry about hallucinations, about mistakes, about a model saying something that isn’t strictly what the company would say, about interactions that don’t feel empathetic enough. 

All fair concerns. But it’s worth holding them up against the alternative. Are human agents perfect? No. They’re people. They make mistakes. They have good days and bad days, and personal or professional pressures that occasionally show up in how they treat a customer. We accept that, manage it, and coach around it. We don’t pretend it never happens and lock it away. 

So what happens when the conversational agent is no longer distinguishable from a person — not by voice, not by latency, not by manner? Does it still make sense to apply a rule that was written for the technology of a previous era?  

That’s a question for regulators and for the industry to work through together, and there are reasonable views on both sides. But there’s a nearer-term consequence that lands squarely on the people who manage these agents, and it’s less about the law and more about control. 

The blind spot we used to be able to ignore 

Here’s the shift.  

When a bot only ever said the exact words we’d hard-coded, there was nothing to check. The greeting was the greeting because it could not physically be anything else. Quality assurance for the automated channel was, effectively, unnecessary. 

Generative AI breaks that.  

A model can express the same intent in its own words. Even when you tell it the greeting should be a certain way, it might phrase it differently — and you won’t know unless you look. The comfortable certainty of “it can only say what we scripted” is gone, and most quality frameworks haven’t caught up. 

Think about how we already treat the human side. Human agents are supervised through quality solutions that measure adherence to scripts, check whether forms are completed correctly, and score different parts of the call. We’ve simply never pointed that lens at the virtual agents, because until recently there was nothing to see. 

That has to change. The same quality form you apply to a human agent now applies to a conversational AI agent — arguably through the very same solution, if you want one unified process across your whole workforce.  

And there’s a bonus in the machine’s favour: because a well-built AI agent should always behave consistently, any deviation from what you expect is a genuine signal. It tells you something has broken, or that customers are using the system in a way you didn’t anticipate. Your QA process stops being a compliance chore and becomes an early-warning system. 

Why this is really a scale problem 

None of this is theoretical by the way. Plenty of organisations get a promising AI pilot working and then stall before it ever reaches full operational impact. The reasons are rarely the model itself — they’re fragmented ownership, thin operational trust, weak frontline adoption, and governance frameworks that were never designed for the messy reality of production.  

In other words: the pilot proved the AI could talk to customers. Nobody built the machinery to know, at scale, how it was talking to them. 

Quality assurance is a big part of that missing machinery. If you can’t see what your virtual agents actually say, you can’t build trust in them, you can’t safely widen their remit, and you certainly can’t defend their behaviour to a regulator, a compliance team or a board. Visibility is the precondition for everything else you want to do with AI in the contact centre. 

So the practical checklist looks less like “buy a model” and more like a governance question. Can you measure what your AI agents say as rigorously as what your people say? Can you spot a drift in tone or an off-script answer before a customer complains? Do your guardrails actually hold, and can you prove it? Is your quality process built for a workforce that is now part human, part machine — and does it treat both to the same standard? 

Answering those questions well is the difference between an AI project that stays trapped in pilot and one that earns its way into production. It’s careful, unglamorous work — designing the guardrails, wiring in the monitoring, unifying quality across human and virtual agents, and building the operational trust that lets you scale with confidence rather than crossed fingers. 

That’s exactly the ground an AI-first expert services partner is built to cover: not just standing an agent up, but making it safe, measurable and ready to grow.  

If you’re trying to move your own AI beyond the pilot and want the quality and governance underneath it to hold, that’s a conversation worth having.