Methodology

How a grade is earned

The discipline of a great inspection guide, applied to phone calls: same tests, blind judges, published evidence — and two separate checks before a grade is allowed out.

Flat engraved emblem: an audio waveform above a row of identical test squares with a small balance scale beside them, inside a double-rule frame
STEP I

We read their promises first

Before a single call, we capture what the vendor claims on its own website and materials — every advertised capability, quoted and dated. The test list is built from those claims plus the industry’s standard capability list, so the report card shows claimed vs. measured, line by line.

STEP II

Controlled calls, run in parallel

Synthetic callers with different voices, accents and dispositions run the same scenario matrix against every vendor — booking, rescheduling, insurance, emergencies, messages. Same tests, same volume, every time. Every call is recorded, and how many calls a grade rests on is printed on the card rather than described in the abstract here.

STEP III

Blind, rubric-bound grading

Judges score each call 0–100 against a written rubric and must cite the exact moments that justify the score. They never know which vendor they’re grading. Measured delays, repeats and dead air feed the grade too.

STEP IV

The failure panel — three labs, majority ruling

Every failing result is re-examined before publication by a panel of three reviewers, each running at a different AI lab, asking one question: was this genuinely the product’s failure — or our caller’s, or a limit of the demo line we called? The majority decides, and a split panel leaves the failure standing. A failure the panel finds was actually correct behaviour is counted as a pass; a limit of the demo line becomes “not testable” and leaves the count; a fault of our own instrument becomes “inconclusive” and leaves it too. Both are disclosed on the card.

STEP V

The second opinion — a judge at another lab

Every judged pass-or-fail is then judged a second time by a model at a different AI lab. Clear agreement settles it. A disagreement — or a score sitting too close to the pass line for agreement to mean much — goes to a person, and nothing resting on that result publishes until they have ruled. Where a person has already ruled on the same criterion, that earlier ruling settles it rather than asking someone the same question twice. Two models agreeing is not a human review, and we do not describe it as one.

These are two different checks, and we keep them straight. The failure panel is three reviewers at three AI labs, voting on failures. The second opinion is one more judge at a different lab, on every judged pass-or-fail, with a person as the tie-break. Neither one is the other, and the second opinion is never “three labs”.

The rules we never break

  • Same test for every vendor. Grades are only comparable because the calls are. Vendors are compared only on identical scenario sets. The scoring criteria are published.
  • Evidence or it didn’t happen. Every published score is backed by the recorded call and the per-line scoring behind it, and a person re-reviews it on any dispute. A verified vendor reads the full transcript of every call behind its own results, the reason we gave each check, and how each result is counted — before the grade even publishes. We do not hand out the audio: our reviewer listens, and what comes back is their written finding, with the result corrected if it was wrong.
  • Honest uncertainty. If the data can’t statistically separate two vendors, we say “indistinguishable” — we never invent a winner.
  • Demo lines are labeled. A grade earned on a public demo says so; demo limitations are disclosed as scope, not counted as failures.
  • Synthetic callers only. Fictional personas, fictional details, explicitly-seeded test records — never real customer data.
  • No lab grades its own family. The caller and the judge run at different AI companies; the failure panel spans three different labs, one per reviewer; and the second opinion always comes from a lab other than the one that judged first. A vendor’s product is never re-judged by reviewers dominated by its own model’s maker.
  • Nothing publishes on one opinion. Two separate checks, not one. Every failing result goes to a panel of three reviewers at three different AI labs, majority ruling. Every judged pass-or-fail also gets a second opinion from a judge at a different AI lab, and a disagreement goes to a person before anything can publish. A grade with a line item still waiting on that person is not published at all — not published late, not published with an asterisk.
  • Claims are quoted, then tested. “Claimed” on a report card means the vendor’s own public materials say so — quoted and dated. A claim we couldn’t verify is marked exactly that, never assumed true or false.
  • Grades age, and say so. Every assessment is a dated snapshot. Vendors are re-tested on a recurring schedule; a new assessment supersedes the old one in the open, and the page always shows when the calls were made.
  • Vendors can always dispute. Free, evidence-based, and public: we either stand by the score with the evidence — or correct it and say we did — see how disputes work.
Want the deeper mechanics — scoring thresholds, statistical method, judge calibration? The engineering write-up is published with every scoreboard.
Proofground

The independent guide to AI receptionists. We make the calls, grade the evidence, and publish it — so you can choose with confidence.

© 2026 Proofground, published by Chest LLC. Grades are based on our own recorded test calls, scored against a published rubric and independently re-reviewed. We take no vendor money for grades — how we stay independent.