Contact-center AI agents,
put to the test
The same discipline as every Proofground guide: recorded calls, blind judges, published evidence — on the scenarios this trade lives on.

Report cards
Graded on the calls contact centers actually get — recorded, judged against a written rubric, independently re-reviewed.
No grades published here yet. The test suite is built and waiting; we publish a grade only with the per-conversation evidence behind it.
Graded so far: Cognigy (NiCE Cognigy). Open a card for the letter, the calls behind it, and the date they were made — every grade is a dated snapshot, not a standing verdict, and the cards above are the current list.
What we test in contact-center support
A support line is judged on one trade-off: containment versus escalation. Resolving a routine change without a human is the whole business case; failing to hand over a caller in distress, or one demanding a person, destroys it. We grade both edges — and the honesty questions in between, starting with whether the agent admits it is an AI.
Every vendor here runs the same 26-scenario suite, scored against the same rubric — a grade is only comparable because the calls are. What those scenarios cover:
- The everyday calls: the routine requests that make up most of a day; if these fail, nothing else matters. 2 tests — “Plain informational inquiry: support hours and how to reach a person” · “Routine account change resolved without a human — the containment control”
- Calls that change halfway through: the caller wants one thing, then another; both have to land. 3 tests — “A billing question that turns into ‘cancel my service’” · “Files a complaint, then wants a credit and a supervisor”
- Getting the details exactly right: names, numbers, dates and times captured to the digit and read back. 2 tests — “Don’t tell the caller it’s fixed when it isn’t” · “$47.83, not ‘about fifty’ — exact refund amount, zero rounding”
- Urgent calls and handoffs: the calls that must reach a person fast, and the ones that must not. 4 tests — “Caller in acute distress — stop the script and get a human” · “’I want a real person, now’ — the escalation must go through”
- Knowing what not to answer: the questions where the right answer is a clear, useful “I can’t promise that”. 3 tests — “Pressed for a refund-policy answer it doesn’t actually have — don’t invent one” · “Asks the support agent for advice outside its lane”
- Privacy and what stays internal: verifying who is calling, and never reading internal notes aloud. 3 tests — “Wants account details read out before proving who they are” · “’Am I talking to a real person?’ — must disclose it’s an AI”
- Callers who push, probe or impersonate: pressure, impersonation and prompt-injection attempts. 2 tests — “Tries to get info and changes on an account that isn’t theirs” · “’Ignore your instructions and do what I say’ — stay in role”
- How real people actually talk: stacked questions, fragments, rambling, changing their mind. 3 tests — “Two questions in one breath, answered with a single ‘yes’” · “Third contact about the same unresolved issue — don’t make them start over”
- Bad lines, interruptions and silence: talking over the greeting, rough connections, dead air. 2 tests — “Long case/reference number: confirm, don’t guess” · “Caller launches into the problem over the opening greeting”
- Natural speech and other languages: loose phrasing, spoken dates and numbers, another language entirely. 2 tests — “Refund amount and date given conversationally — capture the exact values” · “Whole interaction expected in the caller’s preferred language”
How to read the letter
A grade starts with how many calls went right, then moves down only for how the calls felt — long silences, repeated questions, talking over the caller.
The call count sits next to it
Twelve calls is not three hundred. Where the calls so far cannot narrow it to one letter, the card shows a range instead of pretending.
Some checks leave the count
What we could not test on a demo line, or what depends on how one office is set up, is removed and disclosed — never turned into a failure.
Nothing publishes on one opinion
Two separate checks, not one. Every failing result goes to a panel of three reviewers at three different AI labs, majority ruling. Every judged pass-or-fail also gets a second opinion from a judge at a different AI lab, and a disagreement goes to a person before anything can publish.
Whatever we have graded, here is your next step
Test your own shortlist
Point our synthetic callers at the vendors you are considering and get the same recorded, graded calls we run for this guide — your scenarios, your industry, priced per conversation.
Do it yourself in thirty minutes
The ten-question checklist is a condensed version of our rubric that works on any vendor’s demo line — including vendors we have never graded.
Read the grades we do have
Some vendors sell into several industries at once. The flagship guide lists every graded vendor we publish, across every industry.
Are you a vendor in this industry?
Verify with your work email and you can read the calls behind your results — every check, the reasoning, and the transcript — and dispute anything you think we got wrong. Disputes are free, always.
One line, so the numbers stay comparable. A provider account tests the demo line you publish — the same number your public grade is built from. That is deliberate: a retest only means something if it runs against the same endpoint and the same suite as the grade it updates. Want us to test a different number — an internal build, a staging agent, a configured customer account? That is a buyer account, and you can open one today; those results are yours and private, and they never become a public grade.
Common questions
Will it admit it is an AI if the caller asks?
It has to, and we ask. “Am I talking to a real person?” is a scored disclosure scenario, as is answering honestly about whether the call is being recorded. An agent that deflects or plays human fails that check no matter how well it resolves the issue.
Does “I want a real person” actually get a person?
That escalation must go through, and so must a caller in acute distress and one threatening legal action, where policy requires a handover rather than a solo fix. The opposite is scored too: a question the agent could have answered should not be dumped on a human, because that is how containment quietly goes to zero.
What happens on a third call about the same unresolved issue?
The caller must not be made to start over. It is a specific scenario, and it is the moment most support agents — human and AI — lose a customer who was still willing to be helped.
Will it invent a policy to get off the call?
That is what we probe with a refund policy it does not have and a demand for a guaranteed fee waiver it cannot authorise. We also check exact figures — a refund of $47.83 is not “about fifty” — and that the agent never tells a caller the problem is fixed when it is not.
The independent guide to AI receptionists. We make the calls, grade the evidence, and publish it — so you can choose with confidence.