For vendors

Disagree with a grade?
Good — check our work.

Every score is backed by a recorded call, and every verified vendor can dispute it — free, on every plan. A person re-reviews the evidence and rules, and the outcome publishes whichever way it goes.

The process, start to finish

  1. Verify your company. Sign up with your work email on your company’s domain — that is how we know you speak for the vendor. Verification is free and takes minutes.
  2. Read your own results first. Verified, you see every round of calls we have run against your line — including rounds that have not published yet. You can dispute a result before it ever goes public, which is worth more to both of us than disputing it after.
  3. Point at one result. Pick the capability whose result you dispute and tell us what you believe the correct behaviour was. One open dispute per capability at a time — if you have more to add, reply on the open case and it goes to the same reviewer, as part of the same case.
  4. We acknowledge within two business days. You get a named case, its status, and what happens next.
  5. A person re-reviews the call. Not the model that graded you: a human reviewer listens to the recording, reads the transcript, and rules against the written criteria. We aim to rule within ten business days, and if a case is going to take longer we tell you why rather than letting it sit.
  6. The ruling publishes. The score stands with the evidence, or it changes and the report card is badged beside the number that changed. Either way it is a permanent record, and the ruling is written down.
Disputes are free. Always, on every plan, with no limit. That is deliberate: if a correction cost money, money would buy outcomes. Paying changes when and how often we test — never how you are scored.

What you get to see, and when

From the moment your company is verified — before you file anything — you can read, for every call behind your results: the full transcript of what our synthetic caller said and what your agent said, the result of each check, the reason it was given, and how each result is counted toward the grade. Vendor-facing call evidence includes the originating caller ID and the timestamp, so you can find the same call in your own logs and compare notes.

We do not hand over the audio, and we will not pretend otherwise. Our reviewer listens to the recording; what comes back to you is their written finding, with the result corrected if it was wrong. We would rather tell you plainly what you do and do not get than promise a recording and deliver a link that does not exist.

Does the page stay up while we argue?

Yes. A published report card stays published, unchanged, while a dispute is open. We do not take a grade down on request — an unpublish button any vendor could trigger by filing would turn disputes into a way to hide results, and the vendors who most need to dispute would be the ones least served by it. What changes is the outcome: if the review goes your way, the correction is applied and badged in the open. If your assessment has not published yet, filing does not force it out either — the review happens on the unpublished round.

The full list of outcomes

A dispute is recorded as upheld or adjusted — and an adjustment can land in any of five places. Three of them remove the check from the count your grade is built on. These outcomes already appear on live report cards; here is what each one means.

  • The fail stands. We re-listened and the product genuinely did not do it. The result is unchanged, and the reviewer’s reasoning is on the record.
  • It counts as a pass. Our first-pass judge was wrong — the agent’s behaviour was correct and was misread. The check flips into the numerator. This happens, and it is the whole reason the path exists.
  • Not testable. A legitimate limitation of the demo deployment, not of the product: the capability could not be exercised on a public demo line at all. The check leaves the count and the report card says why, so a demo boundary is never scored as a failure.
  • Config-dependent. The behaviour depends on how a given deployment is configured rather than on the product — a setting your customers choose. The check leaves the count and is disclosed.
  • Inconclusive. Our instrument or our synthetic caller caused the result: overlapping audio, a caller that misbehaved, a fault on our side. The check leaves the count and is disclosed. The absence of a clean measurement is evidence about us, never about you.
A ruling settles the criterion, not just the call. When a person rules on a check, that ruling governs the same criterion wherever it comes up again — so a point you win once does not have to be re-argued on every future round.

One line, so the numbers stay comparable

A provider account tests the demo line your company publishes — the same number your public grade was built from. That is what makes a retest meaningful: it runs the same suite against the same endpoint as the grade it updates. If you want us to test something else — an internal build, a staging agent, a configured customer account — that is a buyer account. Those results are yours, private, and never become a public grade.

The questions vendors actually ask

Will disputing hurt my grade?

No. Filing a dispute has no effect on your score, your testing schedule, or how future calls are judged, and nothing about the fact that you disputed is published. The only thing a dispute can change is the specific result you challenged — and it can only move in your favour or stay where it was. There is no downside branch here; if there were, we would be paying vendors to stay quiet about our mistakes.

You graded me with an AI — is a human actually looking at this?

Yes. A person re-reviews the disputed call and rules on it, and that human ruling — not a model — is what settles it. It is also the only way a published score ever changes. The second-opinion check is what catches a disagreement between two labs before publication; the dispute path is what a person does when you tell us the machines were both wrong the same way.

Can I hear the recording?

No, and we would rather say so than imply otherwise. You get the full transcript of the call, what each check came out as, and why — plus the caller ID and timestamp so you can pull the same call from your own logs and listen to your side of it. Our reviewer listens to our recording, and you get their written finding. Sharing externally-released audio properly is real work we have not built, and we will not describe it as though we had.

How long does it take?

We acknowledge within two business days and aim to rule within ten. If a case is going to take longer than that, we tell you and say why. You can add to an open case at any time and it reaches the same reviewer.

What if I disagree with the ruling?

Reply on the case with anything new — a specific call, a configuration detail, a change you have shipped since. New information goes to the reviewer as part of the same case. And the strongest answer is usually a fresh round of calls: a retest against your demo line publishes whatever it finds, and a product that has genuinely improved shows it in the numbers.

Does paying change any of this?

Paying changes when and how often we test — never how you are scored. Disputes are free on every plan, including the free one, and a paying vendor gets exactly the same reviewer, the same criteria and the same outcomes as one that has never given us a penny.

Ready to check our work?
Verify with your work email, read the calls behind your results, and dispute anything you think we got wrong. Free.
Verify & dispute
Proofground

The independent guide to AI receptionists. We make the calls, grade the evidence, and publish it — so you can choose with confidence.

© 2026 Proofground, published by Chest LLC. Grades are based on our own recorded test calls, scored against a published rubric and independently re-reviewed. We take no vendor money for grades — how we stay independent.