Reading a report card

What an A actually means

One scale for every vendor: how often it got the caller’s task done, adjusted by how the calls actually went. Quality can lower a grade. It can never raise one.

The Proofground laurel seal

The letters

One scale, every industry, every vendor. What separates them is how often the agent got the caller’s task done — and how the calls felt while it did.

Last reviewed . This page describes the method in force on that date. When the method changes, this page and this date change with it.

A
Handles nearly everything a great front desk would — reliably, and pleasantly.
B
Solid on the essentials, with specific gaps we name on the report card.
C
Works on the easy path; struggles once a call gets real.
D
Frequent failures on tasks your callers depend on.
F
Did not complete the tests, or failed critical safety moments.

A grade is two things multiplied

Every letter comes from one calculation, used identically on every page of this site:

How often it got the task right, adjusted by how the calls actually went.

The first part is a simple count: of the situations we tested, how many did the agent actually complete correctly? Booked the right slot, moved the right appointment, escalated the emergency, took a message that matched what the caller said. That number is always printed beside the grade, with the number of conversations it came from — you never have to take the letter on faith.

The second part is how the call felt: dead air, repeated questions, talking over the caller, an agent that got there in the end but made the caller work for it. That’s buyer-relevant. An agent that books the appointment while making someone repeat their date of birth four times is not the same product as one that does it cleanly, and pretending otherwise would flatter it.

Quality can only ever lower a grade, never raise one. A charming agent that fails a third of its bookings cannot be talked up into a B. And there’s a floor: call quality can pull the grade down, but never below about seventy percent of the correctness figure — a great task record stays visible even behind a rough experience. Getting the job done sets the ceiling; how it went can only cost you.

Why a grade is sometimes a range

You’ll see grades written as a range — “could be B to D on 12 calls” — instead of a single letter. That is not indecision. It is what the evidence honestly supports.

Twelve calls cannot tell you as much as three hundred. When we’ve run twelve, the true pass rate could reasonably be anywhere across a wide band, and that band can straddle two or three letters. Printing the middle of it as “C” would be inventing a precision we don’t have. So we print the range, and it narrows as we run more conversations. A range is a statement about how much testing has happened, not about how good the vendor is.

You’ll see the same discipline when two vendors are compared: where the difference between them could be the luck of which calls we placed, we say the results are indistinguishable rather than naming a winner.

“Not testable” and “inconclusive” are not failures

Some results don’t belong in a score in either direction, and this is the part of our method we most want you to understand, because getting it wrong means publishing our own problems as a vendor’s.

  • Not testable — the vendor’s demo is deliberately limited in a way that makes the test meaningless. A demo line that never actually writes to a calendar cannot be marked down for not writing to a calendar.
  • Depends on configuration — the answer would be right or wrong depending on how a particular office set the product up, not on the product itself. Whether they accept a specific insurance plan is a setting, not a capability.
  • Inconclusive — something about our own equipment, or our own caller, makes the result unreadable. If we can’t tell whether the agent failed or our equipment did, the answer is not “the agent failed”.

Every one of these leaves the score entirely — it isn’t counted as a pass and it isn’t counted as a fail — and every one of them is disclosed on the report card so you can see how much of the test they accounted for. A vendor is never marked down for something we couldn’t observe. The absence of evidence is a fact about our testing, not about their product.

The reverse also holds: a genuine failure stays a failure. These categories exist to keep us honest, not to give anyone an exit.

Why a vendor can ace the tasks and still score lower

It happens often enough to be worth explaining. A product can be excellent at the mechanics — it books, it reschedules, it escalates — while being unpleasant to talk to: long pauses before each answer, questions asked twice, a tendency to speak over people. The tasks pass; the experience drags the letter down.

That’s the right answer for a buyer. Your callers don’t experience your pass rate. They experience the call.

When there is no grade at all

If nothing scorable came out of an assessment, we publish no grade and say “not yet scored”. We never print an F, and we never print a zero, for a test that didn’t produce a measurement. An F means we measured failure; a blank means we measured nothing.

Similarly, when we have the task record but not enough about how the calls went, the grade is just the task record and the page says the experience wasn’t weighed — rather than implying it was fine.

One honest limitation, stated plainly: the experience half of a grade is scored by AI judges reading the conversation, and every page that shows it labels it as judged. The task half is machine-checked wherever it can be. We tell you which is which on every report card rather than blending them into one number and hoping you don’t ask.

Common questions

Is an A the same in dental as in freight?

The scale is, the tests aren’t. Each industry is graded on the situations that trade actually lives on, so an A means “excellent at this job” in both — but you should compare vendors within an industry, not across.

Why does one vendor have a single letter and another a range?

Number of conversations. A range means we haven’t yet run enough calls to pin a single letter honestly. It narrows as testing continues.

Does a failed emergency call sink the whole grade?

Critical safety moments are weighted heavily and a failure on one is called out on the report card in its own right — precisely so a strong average can’t bury it.

Can I see the tests behind a grade?

Yes. Every report card links to every scored line item and the conversation it came from, and the full method is on the rubric page.

Do grades expire?

They date. Every assessment says when the calls were made, and vendors are re-tested on a rolling schedule; AI products change fast enough that a grade without a date would be a rumour.

Proofground

The independent guide to AI receptionists. We make the calls, grade the evidence, and publish it — so you can choose with confidence.

© 2026 Proofground, published by Chest LLC. Grades are based on our own recorded test calls, scored against a published rubric and independently re-reviewed. We take no vendor money for grades — how we stay independent.