How to evaluate AI agents for customer service

Caroline Adamec
Caroline Adamec
Content Engineer
How to evaluate AI agents for customer service

Key takeaways

  • A demo shows how an AI agent performs on its best day, but testing for genuine resolution, safe failure, and clean human handoff shows how it performs on a normal one.
  • The certifications a vendor holds today are what get a deal through security review.
  • Delight.ai grounds its own numbers in containment and resolution rate rather than a deflection rate, which only measures whether a conversation left the queue, not whether the problem actually got solved.

91% of customer service leaders say they are under pressure to implement AI in 2026, according to Gartner. Most teams start evaluating AI agents for customer service by watching demos, and every demo looks good because a demo is built to succeed, not to reflect the reality of an actual support queue.

Buying an AI agent that resolves customer issues, not just handles a scripted conversation, takes more than one persuasive call with a vendor. 

The finance leader signing the contract needs different proof than the customer service leader who will own the resolution numbers. A small business owner evaluating the same category cares about neither's spreadsheet, just whether the phone gets answered at 9pm on a Saturday. 

This guide walks through how to evaluate AI agents for customer service so everyone in that room, whichever seat they sit in, ends up testing for the same five things.

What is an AI agent for customer service?

An AI agent for customer service is software that understands a customer's request, decides what to do about it, and takes action. That distinguishes it from a scripted chatbot, which matches a customer's words to a pre-written decision tree and breaks the moment a question falls outside that tree.

The difference between a chatbot and an AI agent comes down to reasoning versus keyword matching. 

A chatbot can tell a customer their order shipped because it matched the words in their message to a script, while an AI agent can check the order status, notice it is delayed, rebook the delivery window, and message the customer before they ask, because it reasons through the situation instead of just matching keywords. That shift is exactly why customer service leaders are under so much pressure to evaluate this category well, rather than assume any AI vendor is close enough to any other.

What does it mean to evaluate an AI agent for customer service?

Evaluating an AI agent for customer service means testing it against conditions a sales demo is not designed to show you. A demo runs on curated questions, a clean knowledge base, and a vendor in the room to catch anything unexpected. Production needs to run on your own customers, your live systems, and the edge cases nobody scripted for, with no supervision.

The 5 things a demo won't show you

  1. Whether a resolution is genuine.
  2. Whether the agent fails safely when it does not know an answer.
  3. Whether its compliance posture is current rather than aspirational.
  4. Whether escalation to a person preserves context.
  5. Whether a proof of concept was built to resemble production.

How to evaluate AI agents for customer service

Evaluating AI agents for customer service comes down to five checks, and they work best in order rather than picked at random. A vendor that resolves conversations well but fails silently the moment it does not know an answer will look strong on a scorecard and still damage a customer relationship. The next 5 sections walk through:

  • Genuine resolution
  • Safe failure
  • Current compliance
  • Careful escalation
  • Production-realistic proof of concept

It's worth working through each one before signing, rather than accepting an adjective where a vendor should give you a specific number.

Defining a genuine resolution

A resolution rate is only as honest as its definition. Most customer success teams define a genuine resolution as a conversation where the customer's actual issue was identified, addressed, and confirmed fixed, not just one where the conversation ended. 

Some vendors blur that definition by counting a conversation as resolved the moment a customer does not explicitly ask for a human, instead of requiring the customer to confirm the issue is actually fixed. Those are 2 different numbers wearing the same label, and the gap between them is usually where a disappointing deployment starts.

Ask for the re-contact rate, not just the resolution rate

The sharper question is the re-contact rate, the percentage of "resolved" conversations that bring the same customer back with the same issue within a few days. A high re-contact rate exposes a resolution rate that is really measuring how a customer gave up, not how well their problem got solved.

In its own voice, Delight.ai favors containment and resolution framing over a deflection rate, because deflection only measures whether a conversation left the queue, not whether the customer's problem actually got fixed. 

Part of what makes an honest resolution rate possible is context. A support agent that already knows a customer's order history and past conversations, like Delight.ai's Agent Memory Platform (AMP), does not have to guess at intent the way a stateless chatbot does.

When Hanssem deployed an AI agent to genuinely resolve conversations rather than just deflect them, its resolution rate moved from 48% to roughly 86% over five months, with transfers to a human agent cut in half over the same period.

How do you know an AI agent won't go off-script?

An AI agent that cannot show you why it said what it said is not ready for customer-facing autonomy, however smooth its demo sounds. A platform that hallucinates confidently is worse than one that says it does not know, and your evaluation needs to surface that difference. The National Institute of Standards and Technology's AI Risk Management Framework names four functions every trustworthy AI system needs:

  1. Govern how the system is allowed to behave.
  2. Map its risks to the specific use case.
  3. Measure those risks on an ongoing basis.
  4. Manage the response when something goes wrong.

What to check before you trust an agent with a live conversation

Ask a vendor to show you the reasoning behind a specific answer, not just the answer itself. A platform built for this will have an audit trail that traces a response back to the knowledge source it came from, plus a way to flag a low-confidence answer before a customer ever sees it. 

Delight.ai's governance layer, called Trust OS, gives teams that kind of visibility into why the agent said what it said, along with an emergency switch to pause automated responses without losing any configuration.

Which compliance and security certifications should you require?

The certifications a vendor holds today are what clear your security review. A missing certification does not just slow down procurement, it can disqualify a vendor once security or legal actually reviews the paperwork, and it leaves your business carrying the compliance risk of a partner that has not proven its own controls. At minimum, a customer-facing AI agent should hold:

  • SOC 2 Type II, which confirms security controls operated effectively over a period of time, not just that they were designed correctly at a single point.
  • HIPAA, required if any customer data touches healthcare information.
  • GDPR and CCPA, required for EU and California customer data respectively.

Watch for a Type I report presented as if it carries Type II rigor. A Type I attests that controls are designed correctly at one point in time. A Type II attests that they operated effectively over a period, which is the stronger claim. Delight.ai holds a HIPAA Type I report and a SOC 2 Type II report, along with GDPR and CCPA compliance, and is explicit about which report is which.

How should escalation to a human work?

The moment an AI agent hands a conversation to a person is the moment most deployments either build trust or lose it, and the deciding factor is whether context comes with the handoff. A customer who has to re-explain their issue after being escalated experiences that escalation as a failure, even if the agent's first 90% of the conversation went well.

What should transfer when a case escalates

A clean handoff carries three things to the human agent:

  • The full conversation history, not a one-line summary.
  • Whatever information the AI agent already gathered or confirmed.
  • A plain note on what has already been tried.

Delight.ai's AI-native helpdesk, called Desk, keeps the AI agent and human agents on one shared platform for exactly this reason, so escalation is a handoff within one system instead of a transfer between two. None of this replaces your team. The point of a good AI agent is freeing your team to spend their time on the cases that need a person's judgment, not routing every case to a human by default.

How do you run a demo or POC that predicts production performance?

A proof of concept only tells you something useful if it is built to resemble production. Most POCs run on clean data, a narrow set of questions, and close vendor support, which is exactly why they tend to outperform the deployment that follows them. 

Forrester's 2026 customer service predictions expect service quality to dip industry-wide this year as companies run into exactly this gap between AI's promise and the operational work production actually requires.

Build a POC that resembles production, not a demo

It can be helpful to load the vendor's platform with your own knowledge base and a sample of your real customer queries, rather than a curated set. It can also help to include the requests that break scripted systems alongside the simple ones any AI agent can answer:

  • Multi-step requests, like a refund tied to a delayed shipment.
  • Account changes that require verifying identity first.
  • Exceptions that fall outside your standard policy.

Watch for graceful failure, not just correct answers

A good way to test the AI agent is to ask an ambiguous question on purpose. A good agent will ask a clarifying question or makes a reasonable inference, while a weak one either guesses with confidence or loops the customer in circles. 

That distinction matters because neither the agent, nor your own team, can offer the right solution without first understanding the customer's actual problem, and an agent that fails gracefully asks for the context it's missing instead of guessing past it. How a platform fails when it does not know something tells you more than how it performs when it does.

What else should the buying committee ask?

A handful of criteria round out the evaluation and belong in one shared scorecard instead of five separate vendor conversations.

CriterionWhat to askWhat's at stake
Pricing modelPer seat, per conversation, or per resolution?A per-seat model can cost more at scale than an outcome-based one, even at a lower headline rate.
Implementation timeDays, weeks, or months from contract to first live conversation?Every week in implementation is a week of unresolved conversations and delayed return on investment.
Integration depthCan the agent take real actions in your systems, or only answer questions?An agent that can see an order is late but cannot rebook it is answering, not resolving.
Channel coverageDoes the same agent work across chat, voice, SMS, and email?Customers switch channels mid-conversation, and context should travel with them.
SLA and uptimeWhat is the documented uptime over the last 12 months, and how are outages handled?A vendor that will not share this number usually has a reason not to.

For the engineering side of testing an agent your own team is building, rather than evaluating one to buy, see the complete guide to AI agent evaluation.

The takeaway

The teams that choose well do not get lucky with AI vendors. They ask the same five questions of every platform on their shortlist. Use this as your checklist for evaluating an AI agent for customer service:

  1. What counts as a resolution?
  2. What happens when the agent does not know an answer?
  3. Which certifications are actually current?
  4. What transfers at handoff?
  5. Whether the proof of concept looked anything like production.

Bring this checklist into your next Delight.ai demo, and see how Delight.ai answers each one.

Frequently asked questions