Skip to main content

Omningage

©2026 Omningage. All Right Reserved

If AI is becoming part of the contact centre workforce, it needs more than uptime and containment. It needs performance management. 

We’ve spent decades working out how to measure people in contact centres. 

We listen to calls. Score interactions. Check compliance. Measure customer satisfaction. Look at resolution. Coach people when something goes wrong. Recognise them when they do something well. 

And generally spend quite a lot of time trying to understand whether our agents are doing a good job. 

So who is doing the same thing for the AI? 

Because AI isn’t just sitting quietly in the background anymore. 

  • It’s talking to customers. 
  • Finding information. 
  • Making recommendations. 
  • Accessing systems. 
  • Making decisions. 
  • And increasingly, actually doing things on behalf of the customer and the business. 

The question can’t just be: “Did the AI work?”  It needs to be: “Did the AI do a good job?” 

Working and performing are not the same thing 

Imagine an AI agent handles a customer conversation from beginning to end. 

  • No technical errors. 
  • No transfer to a human. 
  • Customer hangs up. 

Success? Maybe. 

But what if it gave the wrong answer? Followed an old policy? Completed the wrong action? Made the customer repeat themselves four times? Or should have handed over to a person but didn’t? 

The technology worked. The customer experience didn’t. 

That is the difference between measuring whether something operates and measuring whether it performs. 

We already understand this with humans 

Nobody would run a contact centre and say: 

“Our agents answered 98% of their calls without the phone crashing, so everything must be fine.” 

We look deeper. 

  • Did they understand the problem? 
  • Did they give the right answer? 
  • Did they follow the process? 
  • Did they solve it? 
  • Was the customer happy? 
  • Could they have handled it better? 
  • Do they need coaching? 

That is normal contact centre performance management. So why wouldn’t we ask similar questions of AI? 

AI needs QA too 

This isn’t theoretical anymore. AWS is already building tooling that lets contact centres evaluate self-service interactions against business criteria and inspect how AI agents reasoned, used tools, and handled interactions. 

That tells me the conversation is moving from “Is the AI available?” to something much more useful: 

“How well is the AI performing?” 

So what should we measure? 

I don’t think this needs to become complicated. Start with some of the same basic questions we would ask about a person. 

Did it understand? Did the AI correctly understand what the customer actually wanted, or did the customer spend half the conversation trying to correct it? 

Was it accurate? Did it give the right information, based on current knowledge and policy? 

Did it actually solve the problem? Containment on its own is not enough. A customer giving up is technically contained too. 

Did it take the right action? If AI changes an address, makes a booking, updates a record or starts a process, did the right thing happen in the right system for the right customer? 

Did it know when to stop? Sometimes the best thing an AI agent can do is recognise uncertainty, risk or emotion and involve a person. 

Was the handover any good? If it transferred to a human, did the agent get the context, or did the customer have to start again? 

How did the customer feel? Did sentiment improve? Did the customer repeat themselves, abandon or come back again about the same thing? 

The scorecard should change. 

For years, much self-service reporting has centred on containment, deflection, and automation rate. Those numbers still matter, but they do not tell us enough. 

If 70% of customers stay in your automated journey but a chunk of them still don’t get what they need, you’ve created a very efficient way to disappoint people. 

A better AI performance scorecard might include: 

  • Resolution: did we solve the customer’s actual problem? 
  • Accuracy: was the answer correct? 
  • Action: did the right thing happen in the underlying system? 
  • Customer experience: was the journey easy? 
  • Compliance: did the interaction follow the rules? 
  • Handover: did AI know when and how to involve a human? 
  • Improvement: what are we learning from the interactions that do not work? 

Your contact centre increasingly has two workforces 

This is where I think things get really interesting. 

We are beginning to operate contact centres with two types of worker: 

People  +  Digital agents 

The way we improve them is obviously different. 

If a person struggles with a type of interaction, we might coach them. If an AI agent struggles, we might change its instructions, knowledge, tools, guardrails or process. 

But the basic performance loop is not actually that different. 

Evaluate  →  Understand  →  Improve  →  Measure again 

Human: Evaluate → Coach → Improve 

AI: Evaluate → Optimise → Improve 

And both should ultimately be judged against the same thing: 

Did the customer and the business get the right outcome? 

AI performance depends on the climate around it. 

Just like human agents, AI agents do not operate in isolation. Its performance depends on what surrounds it: 

  • Knowledge: is the information reliable and current? 
  • Data: does it have the right customer and interaction context? 
  • Processes: is it being asked to follow something that actually works? 
  • Technology: can it safely access the systems and tools it needs? 
  • Customer journeys: is the AI improving the journey or just sitting inside a bad one? 
  • Leadership and governance: who owns the outcome when it goes wrong? 
  • People: who is responsible for reviewing, improving and challenging it? 

That is why AI quality is not simply an IT job. It is a contact centre performance job. 

You don’t need millions of AI interactions to start 

This does not only matter for businesses running huge autonomous AI programmes. 

If you have introduced one AI-powered customer journey, start looking at it. 

  • Pick a sample. 
  • Find the failures. 
  • Look at what customers repeat. 
  • Look at where they abandon. 
  • Look at the transfers. 
  • Look at what happened after the interaction. 
  • Make one improvement. 
  • Then measure it again. 

The same principle applies whether you are running 100 automated conversations or 10 million. 

Small improvements still count. 

The question is going to become difficult to ignore 

Imagine that a few years from now, 40% of customer interactions are handled by people and 60% are handled primarily by AI. 

But 95% of the quality management effort is still focused on humans. 

That would seem pretty odd. 

Yet I suspect plenty of organisations are heading in that direction. 

Who is quality-assuring your AI agent? 

Not just checking whether it is online. Not just measuring how many contacts it contains. 

Actually understanding: Did it do the right thing? Did it deliver a good customer outcome? And is it getting better? 

Because if AI is going to become part of the workforce, it should probably be expected to perform like one. 

Context: AWS introduced automated evaluation for self-service interactions and AI-agent trace tooling in Amazon Connect in 2026. 

 

×
×
×