The rubric, in public
The list is written before any calls are placed, every vendor in an industry runs the identical one, and every published score links to the conversation it came from.

Last reviewed . This page describes the method in force on that date. When the method changes, this page and this date change with it.
What a rubric is here
A rubric is the written list of things a conversation is scored on, fixed before any calls are placed. Every vendor in an industry is scored against the same list, in the same order, on the same situations. That is the only reason two grades on this site can be compared at all.
Below is the structure — the kinds of line items, how each kind is decided, and the rules the list itself has to obey. Individual industry rubrics carry the specific wording, and every report card links to the exact line items behind each score, with the conversation that produced them.
The kinds of things we score
Across industries the list is built from the same families of criteria:
- Task completion. Did the caller get what they called for? Booked, rescheduled, cancelled, routed, quoted, dispatched — the specific outcome that trade’s callers want.
- Accuracy of capture. Name, date of birth, phone number, address, the reason for the call — recorded as the caller actually said them, not as the agent guessed.
- Verification before change. Nothing about an existing record should move until the agent has established who it’s talking to. Scored separately from whether the change itself worked.
- Escalation. Emergencies, distress, and anything outside the agent’s competence reaching a human — or a real emergency path — instead of a calendar slot. In several industries this is a critical item: failing it fails the call outright regardless of everything else.
- Message-taking. What reaches the business matches what the caller said. A message that mangles a name or a callback number creates work rather than saving it.
- Honesty about limits. An agent that doesn’t know should say so and hand off. Confident improvisation — inventing coverage, prices, availability or policy — is scored as a failure, not as helpfulness.
- How the call went. Measured, not judged by feel: time to first response, silences, repeated questions, interruptions, whether the caller had to say things twice.
How each item is decided
Line items fall into three kinds, and which kind an item is determines who decides it:
- Measured items are machine-checked. Did the appointment actually get created; how long the pauses were; whether the call was answered at all. There is no reading to misinterpret, so these are never sent to a judge.
- Capture items check a specific value against what the caller actually said — and check that the agent’s own words in the conversation show it captured that value. If the value happens to match but nothing in the conversation evidences the capture, the honest result is “we could not observe this”, and it leaves the score rather than becoming a failure.
- Judged items are the ones that need a reading of the conversation — whether an escalation was handled properly, whether an answer was honest, whether the caller was treated well. A judge scores the item 0 to 100 against the written rubric for that item and must cite the exact moments in the conversation that justify the score. The score is then resolved to a plain pass or fail at that item’s stated line; the default line is 75 out of 100, and an item whose reasoning can’t be parsed fails closed rather than being waved through.
Judges are blind, and failures get a second lab
- Blind. The judge scoring a conversation does not know which vendor produced it. No name, no branding, no line identity. It cannot favour a customer it cannot identify.
- Not the same family. The AI model playing our caller and the AI model judging the conversation come from different AI labs on purpose, so no model is grading its own performance.
- Failing items are re-reviewed before publication by a review panel drawn from a different AI lab than the first judge, asking one question: was this genuinely the product’s failure — or our caller’s, or a limitation of the demo we were calling? A first-pass fail that the panel finds was actually correct behaviour is counted as a pass; a demo limitation becomes “not testable”; an artefact of our own equipment becomes “inconclusive”. The raw verdict is never overwritten — what changes is how the row counts, and the report card shows both.
- Judged results near the line get a second opinion from a judge at a different AI lab. When the two agree clearly, the result stands. When they disagree, a person rules on it — and nothing publishes until they have. A grade with an unresolved line item is not published at all.
- One honest residual, stated deliberately: two AI models can be wrong in the same way, and agreement between them cannot see its own shared blind spot. That’s why disputes are free, why every score links to its evidence, and why a result resolved by agreement rather than by a person says so.
Rules the rubric itself has to obey
A rubric can be unfair by asking for the wrong things, so the list is written to match what an experienced office manager would actually expect from a good front desk — not what a spec sheet could theoretically demand:
- We never require any-requested-time availability. A good receptionist offers the next real slot; being unable to conjure Tuesday at 4 is not a failure.
- We never require a price quote over the phone. Plenty of excellent front desks correctly decline to quote, and an agent that invents a number is doing the worse thing.
- We never require a per-office capability. Which insurers a practice takes, which languages it speaks, what a particular procedure costs there — those are settings of one business, not properties of the product.
- Our caller follows the agent’s lead. Our synthetic callers behave like real ones: they answer what they’re asked and don’t helpfully front-load information the agent never requested. We grade the agent’s actual flow, not its performance against a cooperative script.
- A rubric version is pinned to its results. When we change a rubric, results from the old version aren’t silently compared to the new one — cross-version comparisons are refused rather than fudged.
And if you think we got one wrong
Every score links to the conversation it came from. Any verified vendor can dispute any result, free, and we either stand by the score with the evidence or correct it and badge the report card as corrected. How disputes work.
Common questions
Can I see the exact rubric text for my industry?
Every report card links to the individual line items behind each score, with the conversation each one came from. That is the rubric as applied — the version of it that actually decided a grade, rather than a marketing summary of it.
Who writes the rubrics?
We do, before anyone is graded on them, from the standard capability list for that trade plus the vendor’s own advertised promises — captured and dated from their public materials, so the report card can show claimed against measured.
Do vendors get to change the rubric?
No. A vendor can dispute how a rubric was applied to a specific conversation, and that is free. Nobody buys their way out of a criterion they score badly on.
Is a human involved at all?
Yes, at the points where it matters: when the two independent judges disagree, or agree too close to the line to be sure, a person rules on it — and the grade does not publish until they have.
What happens to a test you can’t interpret?
It leaves the score and is disclosed. If we can’t tell whether the agent failed or our own equipment did, we will not report it as the agent failing.
The independent guide to AI receptionists. We make the calls, grade the evidence, and publish it — so you can choose with confidence.