You pay for conversations.
That’s the whole meter.
A fixed amount each month plus an amount per conversation, billed at the end of the month for what actually ran. If our own equipment broke the call, you don’t pay for it.
If you’re buying — testing your own shortlist
You bring the vendors you’re weighing and the situations that matter to your business; we place the conversations and grade them. These are the current plans, served live from our billing system — not a screenshot.
Free sample
- Free — no charge
- Our starter set of test situations
- One vendor per test
- Up to 3 calls per situation
Shortlist
- First 100 · $2.50 each
- 101–500 · $1.75 each
- 501 and above · $1.25 each
- Write your own test scenarios
- Build your own caller profiles
- Compare two vendors head to head
- The full report, call by call
- Download your results to share
- See how a vendor compares with its industry
- Our starter set of test situations
- Up to 4 vendors in one test
- Up to 10 calls per situation
Monitoring
- First 300 · $1.50 each
- 301 and above · $1.25 each
- Write your own test scenarios
- Build your own caller profiles
- Repeat the same tests on a schedule
- Compare two vendors head to head
- The full report, call by call
- Download your results to share
- See how a vendor compares with its industry
- Our starter set of test situations
- Up to 8 vendors in one test
- Up to 20 calls per situation
Each band prices only the conversations inside it — the rate steps down as you use more, and it is never applied backwards.
- Every grade cites the exact moments in the conversation it came from — with the recording wherever a call was recorded
- A conversation our platform broke never counts against anyone and is never billed
- When two vendors are too close to call, we say so instead of naming a winner
What the buyer product actually involves — what you need to have ready, and what lands in your inbox — is on Test your shortlist.
If you’re a vendor we grade
Disputing a published result is free and always will be. What a vendor can buy is testing of its own demo line — more of it, sooner.
Retest
- First 200 · $2 each
- 201 and above · $1.50 each
- File disputes on any result
- Request an off-schedule retest of your demo line
Continuous
- First 1,200 · $1.50 each
- 1,201 and above · $1.25 each
- File disputes on any result
- Request an off-schedule retest of your demo line
- Continuous monitoring between scheduled tests
- Download the fix brief for your published results
Each band prices only the conversations inside it — the rate steps down as you use more, and it is never applied backwards.
- Every grade cites the exact moments in the conversation it came from — with the recording wherever a call was recorded
- A conversation our platform broke never counts against anyone and is never billed
- When two vendors are too close to call, we say so instead of naming a winner
How the meter works
One shape, both sides of the market: a fixed amount each month, plus an amount for each conversation you run. Nothing is charged up front — we bill at the end of the month for what actually happened during it.
The per-conversation rate is graduated: it steps down as your volume rises, and each band prices only the conversations that fall inside it. If the rate drops after the first thousand conversations, your first thousand still price at the first rate and only the ones above it price at the lower one. That’s deliberate. The alternative — one cheaper rate applied to everything once you cross a line — produces a bill that goes down when your usage goes up, and a bill nobody can check is a bill nobody should sign.
How many conversations does an evaluation actually take?
This is the question the plan cards can’t answer for you, and without it the rate above is unusable. A run is three numbers multiplied together — how many vendors, how many situations, how many times each — so here are the sizes real evaluations come out at:
- A quick first look: 2 vendors × 5 situations × 3 repeats = 30 conversations. Enough to see which one falls over, and honest about the rest — at that size a grade usually publishes as a range rather than a single letter.
- A typical shortlist run: 3 vendors × 8 situations × 5 repeats = 120 conversations. This is where most buyers start, and it is normally enough to separate a good agent from a bad one.
- A decision you’re putting a year of budget behind: 3 vendors × 12 situations × 10 repeats = 360 conversations. This is the size at which two close vendors usually stop being “too close to call”.
The repeats are not padding. One call tells you what happened once; five tell you whether it happens reliably. And the number of repeats is exactly what decides whether we can honestly name a winner at the end or have to tell you the two vendors are indistinguishable — which is why we’d rather you ran eight situations five times than forty situations once.
Multiply your own three numbers, then read the rate off the plan cards above. We deliberately don’t print a total on this page: the plans are served live from our billing system, and a number typed into the page would drift from the one on your invoice.
Working out what the vendors will charge you is a different exercise, and there’s a page for it: how much an AI receptionist costs.
What counts as a billable conversation
One conversation with a vendor’s agent — one phone call, one chat session, one email thread, one SMS exchange — that ran to a real ending and produced a result we could score. That’s the unit, and it’s the same unit whether the result was good or bad for the vendor: a conversation where the agent failed the task still cost us a call and still tells you something, so it bills.
Three things are not conversations and never bill: a run we place for our own public grading, an attempt that never reached the vendor at all, and — the one worth reading twice below — anything our own platform broke.
If our own equipment broke the call, you don’t pay for it
Every test run depends on our own equipment: the phone system, the synthetic caller, the connection that carries the audio. When that equipment is what failed, the conversation is marked as our fault, it is excluded from the results, and it is never billed to anyone — not to the buyer who commissioned it, not to the vendor being tested.
The same branch in our code does both jobs, and that is the whole point. A platform fault that quietly billed a customer would be a platform fault that quietly counted against a vendor’s grade. We refuse to charge for our own mistakes for exactly the reason we refuse to publish them as somebody else’s.
A card before the first billable test
We ask for a card on file before the first conversation that can bill — even when the fixed monthly amount on your plan is zero. Billing in arrears without a card means discovering at month end that a month of testing was never collectable, and the version of that story where we chase you for it is worse for both of us than the version where we ask first.
Some plans carry a monthly conversation cap. Where one does, we refuse to start a run that would exceed it, before any calls are placed — you are never billed for an overage you couldn’t see coming.
Changing or cancelling
Cancel whenever you like; it takes effect on your plan, not on your history. Conversations that already ran still bill. You’ll receive one final invoice covering the tests placed before you cancelled, and nothing after it.
The price you agreed to when a run started is the price that run is billed at. If we publish new plan pricing while your test is in flight, your test keeps the tariff you approved — a new number can never reach an invoice without having passed a preview and your explicit go-ahead first.
Buying an account means agreeing to the terms of service and the privacy policy of Chest LLC, which publishes Proofground.
Common questions
Is reading the grades free?
Yes, and it always will be. Every published report card, every industry guide and every piece of evidence behind a grade is free to read, with no account and no email required. Money only enters when you ask us to run tests of your own, or when a vendor asks us to test its own line more often.
Why per conversation instead of per minute?
Because a conversation is the thing you actually care about — one caller trying to get one thing done. Per-minute pricing quietly rewards a slow agent and punishes a fast one, which is the opposite of what we are here to measure.
What if a test conversation fails because of your equipment?
It is excluded from the results and billed to nobody. If our own equipment broke the call, you don’t pay for it — and the vendor doesn’t wear it either.
Do I get charged before I see anything?
No. Billing is monthly in arrears, so the first invoice arrives after the first month of testing, covering conversations that actually ran. A card is required before the first billable test starts, but it is not charged at that moment.
Does paying change a published grade?
No. Paying changes when and how often we test — never how you are scored. Disputes, which are the only route to a correction, are free for exactly that reason.
The independent guide to AI receptionists. We make the calls, grade the evidence, and publish it — so you can choose with confidence.