About

We grade named companies.
So here’s our own file.

What we test, how we score it, and exactly where the money comes from — because a publication that grades other people’s products owes its readers that before it owes them anything else.

Last reviewed . This page describes the method in force on that date. When the method changes, this page and this date change with it.

What we do

Proofground is an independent testing outfit for AI agents that talk to customers. We place hundreds of controlled, recorded conversations against each vendor’s deployed agent — phone, chat, email, SMS — score them against a rubric written before the first call, and publish the result with the evidence attached.

The model is borrowed, deliberately, from the inspection guides that already work: someone independent has to actually walk in, order the meal, and write down what happened. Software buyers have never had that. They have had vendor demos, affiliate lists, and review sites collecting opinions from people who mostly haven’t used the product under pressure.

How we’re funded

We grade named companies in public. That obliges us to tell you exactly where the money comes from, so here it is, in full:

  • Buyers pay to run their own tests. Businesses evaluating a shortlist pay us to test their candidate vendors against their situations, and they get that report privately. That is our primary revenue.
  • Vendors pay for cadence. A verified vendor can buy testing of its own published demo line — more often, and sooner, than our public calendar would reach it — along with a private fix brief. Paying changes when and how often we test — never how you are scored. A purchased retest publishes whatever it finds, including a worse grade than the one before it.

And the list of things we do not take, which is the more useful list:

  • No affiliate commissions. Not a cent, from any vendor, on any signup that originates here. A ranking funded by referral fees is not a ranking; it is a commission schedule with opinions attached.
  • No sponsorships. Nobody underwrites a guide, an industry page or a report card.
  • No paid placement. Position in the scoreboard is a consequence of the grade and nothing else. There is no promoted row and no premium listing to buy.
  • No fee to be listed, and none to be removed. Being graded costs a vendor nothing, and a published assessment cannot be bought off the site.
  • No payment for a grade, a delay, or a scenario exclusion. The vendor pricing page lists what money can never buy, item by item.

The masthead: who is accountable for what we publish

Proofground is a service of Chest LLC, which is responsible for everything published here — the grades, the guides, the comparisons and the software behind them. The company was founded by Reza Alavi, MD, MHS, MBA, a physician and adjunct faculty at the Johns Hopkins University School of Medicine.

The work is done by a team, and it is divided into four functions. We describe them because a reader should know which part of the company has to stand behind a claim, and what each one is not allowed to do:

  • Editorial decides what publishes. Nothing reaches the site without its sign-off, and that authority runs in one direction only: it can withhold or correct a result, and it cannot improve one. No commercial conversation, with a buyer or a vendor, reaches this function before publication.
  • The review board writes each industry’s rubric before any vendor is graded against it, and rules on the judged results our two independent AI judges disagree about. A grade with an open disagreement on any line item does not publish until the board has ruled on it — not as a courtesy, as a hard gate in the software.
  • The testing desk owns the calling equipment and the standing instruction that governs it: when results look bad, the first suspect is our own equipment, not the vendor. It is this desk’s job to prove the equipment was working before any failure is attributed to a named company.
  • Corrections are handled in the open. A corrected report card is badged as corrected and keeps the record of what changed. We do not quietly edit a published grade, and we do not remove one.

These are functions rather than job titles, because what protects a reader is which part of the company has to stand behind a claim and what it is forbidden from doing. Nobody with a relationship to a vendor takes any part in grading that vendor — where such a relationship exists anywhere in our work, that company is withheld from publication entirely rather than graded by someone connected to it. And no vendor we grade holds any stake in Chest LLC or any say in what we test, how it is scored, or what publishes.

Standards we hold ourselves to

  • Judges are blind to vendor identity, and the model that plays our caller comes from a different AI lab than the one that judges the conversation.
  • Evidence or it didn’t happen. Every published score links to the conversation behind it, with the moments that produced the score cited in it.
  • We only test lines vendors publish. Never a real business’s production number, and never a customer’s deployment without that customer commissioning it.
  • Synthetic callers only. Fictional personas, fictional details, explicitly-marked test records. No real customer or patient data is ever involved.
  • We distinguish “failed” from “we couldn’t measure it”. When our own equipment is what broke, that is a fact about us — it leaves the score and is disclosed, never published as a vendor’s failure. How that works.
  • Assessments are dated snapshots. They are superseded in the open, never quietly edited, and a correction is badged as a correction.
  • Disputes are free, forever. If challenging a score cost money, we would have an incentive to publish scores worth challenging.

When we get it wrong

We will. Placing thousands of phone calls reliably is hard, and the failure we’ve had to engineer against hardest is our own equipment breaking a call and the result reading like a vendor’s failure. Our standing rule when results look bad is that the first suspect is our own equipment — in practice it has been our equipment far more often than the vendor.

When a published result turns out to be wrong, we correct it, badge the report card, and leave the record of the correction in place. That is the whole deal we’re offering readers: not that we’re never wrong, but that you can check, and that being right is worth more to us than any single grade.

If you’ve found something wrong on this site: vendors should file a dispute, which is free and gets a re-review against the recording. Everyone else can simply write to hello@proofground.ai — a reader who can show us we misread a conversation is doing the most valuable thing anyone does here.

Common questions

Are you owned by, or affiliated with, any vendor you grade?

No. No vendor we grade holds any ownership stake in Proofground, and none has any say in what we test, how it is scored, or what publishes. Nor is any vendor graded by someone with a relationship to it: where such a relationship exists anywhere in our work, that company is withheld from publication entirely rather than graded. On top of that, judges never see vendor identity, every vendor in an industry runs the identical rubric, and no amount of money buys, delays or removes a grade.

How do you choose who to test?

We test vendors that publish a demo line anyone can call, starting with the ones buyers ask us about most. Vendors can request testing of their own line; that sets the schedule, never the result.

Is reading the grades free?

Always. No account, no email, no paywall. Money enters only when someone commissions testing of their own.

Do you take advertising?

No advertising, no sponsorships, no affiliate commissions and no paid placement. The only money that reaches us is payment for testing.

How can I check any of this?

Open a report card. Every score links to the conversation it came from, and anything we could not measure is disclosed rather than scored. The full method is on the rubric page.

Proofground

The independent guide to AI receptionists. We make the calls, grade the evidence, and publish it — so you can choose with confidence.

© 2026 Proofground, published by Chest LLC. Grades are based on our own recorded test calls, scored against a published rubric and independently re-reviewed. We take no vendor money for grades — how we stay independent.