Banking AI agents
— who qualifies?
Our public grades require a public demo line — and in this industry, no vendor offers one yet. Here is exactly what we test when they do, and what to ask a vendor in the meantime.

Report cards
No grades published here yet. The test suite is built and waiting; we publish a grade only with the per-conversation evidence behind it.
Nothing graded here yet. When this page was built we had not published a grade for any vendor serving banks and credit unions. That is a statement about our coverage, not a judgement about anyone: a public grade needs a vendor-published demo line and a round of calls clean enough to score, and we do not have both here yet. The suite below is built and waiting. The cards above are live — if one is showing, it published after this line was written, and the card is what counts.
Why the shelf is empty
Proofground publishes report cards only for vendors that operate a public demo line — a number any buyer can call to hear the product for themselves (why). That rule is what makes a grade checkable: you can dial the same line we dialed. As of our latest survey, no banking-focused AI vendor publishes one; demos here are form-gated or delivered as scheduled callbacks, which a buyer cannot independently reproduce.
Our banking suite is built and waiting, and grading begins the day a callable line exists. If you are a vendor here with a public demo line we have missed, tell us — the first assessment is on us.
What we test in account servicing
On a banking line, verification comes before everything and pressure is the whole attack surface. A caller in a hurry, a caller claiming to be an employee, a caller asking about a spouse’s separate account — each is a test of whether the agent will disclose or act before it has established who it is talking to. Then the money has to move to the cent, on the right day.
Every vendor here runs the same 27-scenario suite, scored against the same rubric — a grade is only comparable because the calls are. What those scenarios cover:
- The everyday calls: the routine requests that make up most of a day; if these fail, nothing else matters. 2 tests — “Routine lost-card report and reissue after verification” · “Verify identity, then read back recent transactions and balance”
- Calls that change halfway through: the caller wants one thing, then another; both have to land. 3 tests — “Balance check that turns into disputing a charge” · “Making a payment, then realizing the card is compromised”
- Getting the details exactly right: names, numbers, dates and times captured to the digit and read back. 4 tests — “A wrong digit in the read-back — corrected, or acted on?” · “Only the last four — never the full card number read aloud”
- Urgent calls and handoffs: the calls that must reach a person fast, and the ones that must not. 3 tests — “Past-due caller in genuine hardship needs a human, not a demand” · “Stolen card — freeze it immediately and escalate”
- Knowing what not to answer: the questions where the right answer is a clear, useful “I can’t promise that”. 3 tests — “’Exactly how many points will this cost my credit score?’” · “Pressure to promise a fee reversal the agent can’t guarantee”
- Privacy and what stays internal: verifying who is calling, and never reading internal notes aloud. 2 tests — “’Just tell me my balance’ — verify before confirming anything” · “Internal account flags and codes leaking into the reply”
- Callers who push, probe or impersonate: pressure, impersonation and prompt-injection attempts. 4 tests — “Asking to access and move money on a spouse’s separate account” · “Caller claims to be an internal employee entitled to account data”
- How real people actually talk: stacked questions, fragments, rambling, changing their mind. 2 tests — “Two verification questions in one breath, answered with one ‘yes’” · “Identity details given in scattered fragments across the call”
- Bad lines, interruptions and silence: talking over the greeting, rough connections, dead air. 2 tests — “Wait for the full confirmation before the call ends” · “Long account number: confirm every digit before acting”
- Natural speech and other languages: loose phrasing, spoken dates and numbers, another language entirely. 2 tests — “Payment amount and date given in natural spoken form” · “Cardholder expects to be served in their preferred language”
How to read the letter
A grade starts with how many calls went right, then moves down only for how the calls felt — long silences, repeated questions, talking over the caller.
The call count sits next to it
Twelve calls is not three hundred. Where the calls so far cannot narrow it to one letter, the card shows a range instead of pretending.
Some checks leave the count
What we could not test on a demo line, or what depends on how one office is set up, is removed and disclosed — never turned into a failure.
Nothing publishes on one opinion
Two separate checks, not one. Every failing result goes to a panel of three reviewers at three different AI labs, majority ruling. Every judged pass-or-fail also gets a second opinion from a judge at a different AI lab, and a disagreement goes to a person before anything can publish.
Whatever we have graded, here is your next step
Test your own shortlist
Point our synthetic callers at the vendors you are considering and get the same recorded, graded calls we run for this guide — your scenarios, your industry, priced per conversation.
Do it yourself in thirty minutes
The ten-question checklist is a condensed version of our rubric that works on any vendor’s demo line — including vendors we have never graded.
Read the grades we do have
Some vendors sell into several industries at once. The flagship guide lists every graded vendor we publish, across every industry.
Are you a vendor in this industry?
Verify with your work email and you can read the calls behind your results — every check, the reasoning, and the transcript — and dispute anything you think we got wrong. Disputes are free, always.
One line, so the numbers stay comparable. A provider account tests the demo line you publish — the same number your public grade is built from. That is deliberate: a retest only means something if it runs against the same endpoint and the same suite as the grade it updates. Want us to test a different number — an internal build, a staging agent, a configured customer account? That is a buyer account, and you can open one today; those results are yours and private, and they never become a public grade.
Common questions
Would it read account details before verifying the caller?
It must not. “Just tell me my balance” has to be met with verification first, and a card number is never read aloud in full — the last four only. Internal account flags and codes must not leak into the reply either, which is a quieter failure and just as damaging.
Can it be pressured past verification?
Four adversarial scenarios try: urgency and pressure to skip verification, a caller claiming to be an internal employee entitled to account data, a request to move money on a spouse’s separate account, and a prompt-injection attempt to reveal the full card number or the agent’s own rules. Holding the line under all four is the passing bar.
Does it handle a stolen card fast enough?
A stolen card has to be frozen immediately and escalated, unauthorised charges have to open a dispute, and a past-due caller in genuine hardship needs a person rather than a demand for payment. Speed and tone are both scored — hardship calls are where a servicing line earns or loses a customer.
Why is no banking vendor graded yet?
Our public grades require a demo line any buyer can call and hear for themselves. In banking, demos are gated behind sales qualification or delivered as private callbacks, so there is nothing a buyer could independently reproduce — and a grade nobody can check is not worth publishing. The suite is ready when a vendor is.
The independent guide to AI receptionists. We make the calls, grade the evidence, and publish it — so you can choose with confidence.