The scoreboard

Every vendor we’ve graded

One list, every industry, dated. Each entry opens onto the scored line items and the recorded conversations the grade was built from.

The Proofground laurel seal

Every graded vendor

All industries, newest assessments first. Every entry links to the full report card — the scored line items and the conversations behind them.

No grades published here yet. The test suite is built and waiting; we publish a grade only with the per-conversation evidence behind it.

How to read an entry

Each row carries the few facts that decide whether a grade is worth anything to you. Read them in this order:

  • The grade. How often the agent got the caller’s task done, adjusted by how the calls actually went — quality can lower a letter, never raise it. A grade written as a range means we haven’t run enough conversations to pin a single letter honestly, and it will narrow as we run more. The full explanation.
  • The industry. Grades are comparable within an industry, because within an industry every vendor runs the identical situations. A dental A and a freight A both mean “excellent at this job” — they are not two scores on one scale.
  • When it was tested. Every assessment is a dated snapshot. AI products change fast; a six-month-old grade is a fact about six months ago, and a newer assessment supersedes it in the open rather than quietly replacing it.
  • How many conversations. The single best guide to how much weight a grade can carry. A few dozen conversations shortlists; a few hundred decides.
  • Whether a critical scenario failed. Emergency escalation, safety boundaries, verification before changing a record. These are called out on the report card in their own right, so a strong average can’t bury one — a vendor can grade respectably overall and still be disqualified by the one call you can’t afford to lose.
  • What it was tested against. A public demo line is what nearly every published grade here is built on: the number the vendor publishes for anyone to call. It shortlists well, but a demo can be more polished — or deliberately more limited — than the product you would deploy. A grade earned against a live account is testing of a real deployment, commissioned by the business that runs it.
Read every grade as a signal, not a verdict. A demo-line grade is a fast, free way to narrow five vendors to two. To know which of those two fits your business, open a pilot with both and run your own situations against the real product — that’s what a private run is for.

Missing a vendor?

Two reasons a vendor you’re considering isn’t here. Either we haven’t reached them yet — tell us who you’re weighing and they move up the calendar — or they don’t publish a demo line anyone can call. We only test lines a vendor has published for public use, never a real business’s production number, so a vendor whose demo is form-gated or delivered as a callback can’t be graded publicly until that changes.

Vendors can also ask to be tested. That sets the schedule and nothing else: paying changes when and how often we test — never how you are scored. Vendors start here.

Narrowed it to two?
Run the same recorded, graded conversations against both vendors’ real pilot accounts — your situations, your call mix, evidence attached.
Test your shortlist
Proofground

The independent guide to AI receptionists. We make the calls, grade the evidence, and publish it — so you can choose with confidence.

© 2026 Proofground, published by Chest LLC. Grades are based on our own recorded test calls, scored against a published rubric and independently re-reviewed. We take no vendor money for grades — how we stay independent.