How a grade is earned
The discipline of a great inspection guide, applied to phone calls: same tests, blind judges, published evidence — and two separate checks before a grade is allowed out.

We read their promises first
Before a single call, we capture what the vendor claims on its own website and materials — every advertised capability, quoted and dated. The test list is built from those claims plus the industry’s standard capability list, so the report card shows claimed vs. measured, line by line.
Controlled calls, run in parallel
Synthetic callers with different voices, accents and dispositions run the same scenario matrix against every vendor — booking, rescheduling, insurance, emergencies, messages. Same tests, same volume, every time. Every call is recorded, and how many calls a grade rests on is printed on the card rather than described in the abstract here.
Blind, rubric-bound grading
Judges score each call 0–100 against a written rubric and must cite the exact moments that justify the score. They never know which vendor they’re grading. Measured delays, repeats and dead air feed the grade too.
The failure panel — three labs, majority ruling
Every failing result is re-examined before publication by a panel of three reviewers, each running at a different AI lab, asking one question: was this genuinely the product’s failure — or our caller’s, or a limit of the demo line we called? The majority decides, and a split panel leaves the failure standing. A failure the panel finds was actually correct behaviour is counted as a pass; a limit of the demo line becomes “not testable” and leaves the count; a fault of our own instrument becomes “inconclusive” and leaves it too. Both are disclosed on the card.
The second opinion — a judge at another lab
Every judged pass-or-fail is then judged a second time by a model at a different AI lab. Clear agreement settles it. A disagreement — or a score sitting too close to the pass line for agreement to mean much — goes to a person, and nothing resting on that result publishes until they have ruled. Where a person has already ruled on the same criterion, that earlier ruling settles it rather than asking someone the same question twice. Two models agreeing is not a human review, and we do not describe it as one.
The rules we never break
- Same test for every vendor. Grades are only comparable because the calls are. Vendors are compared only on identical scenario sets. The scoring criteria are published.
- Evidence or it didn’t happen. Every published score is backed by the recorded call and the per-line scoring behind it, and a person re-reviews it on any dispute. A verified vendor reads the full transcript of every call behind its own results, the reason we gave each check, and how each result is counted — before the grade even publishes. We do not hand out the audio: our reviewer listens, and what comes back is their written finding, with the result corrected if it was wrong.
- Honest uncertainty. If the data can’t statistically separate two vendors, we say “indistinguishable” — we never invent a winner.
- Demo lines are labeled. A grade earned on a public demo says so; demo limitations are disclosed as scope, not counted as failures.
- Synthetic callers only. Fictional personas, fictional details, explicitly-seeded test records — never real customer data.
- No lab grades its own family. The caller and the judge run at different AI companies; the failure panel spans three different labs, one per reviewer; and the second opinion always comes from a lab other than the one that judged first. A vendor’s product is never re-judged by reviewers dominated by its own model’s maker.
- Nothing publishes on one opinion. Two separate checks, not one. Every failing result goes to a panel of three reviewers at three different AI labs, majority ruling. Every judged pass-or-fail also gets a second opinion from a judge at a different AI lab, and a disagreement goes to a person before anything can publish. A grade with a line item still waiting on that person is not published at all — not published late, not published with an asterisk.
- Claims are quoted, then tested. “Claimed” on a report card means the vendor’s own public materials say so — quoted and dated. A claim we couldn’t verify is marked exactly that, never assumed true or false.
- Grades age, and say so. Every assessment is a dated snapshot. Vendors are re-tested on a recurring schedule; a new assessment supersedes the old one in the open, and the page always shows when the calls were made.
- Vendors can always dispute. Free, evidence-based, and public: we either stand by the score with the evidence — or correct it and say we did — see how disputes work.
The independent guide to AI receptionists. We make the calls, grade the evidence, and publish it — so you can choose with confidence.