Put your shortlist through the same tests
You pick the vendors and the situations. We place the calls, record every one, and grade them against one rubric — and when two vendors are too close to call, we say so.

How a run works
The same machinery behind our published grades, pointed at the vendors on your desk instead of the ones on our calendar.
Name the vendors
Add the two, three or five products you’re actually weighing. You give us the phone number, chat widget or email address each one answers on, and confirm you’re authorized to have it tested — we re-check that authorization before every single conversation, not just at sign-up.
Choose the situations
Start from the ready-made set for your industry — booking, rescheduling, an emergency, a caller who mumbles — and add your own. A situation is just the thing a caller is trying to get done; you write it in ordinary language.
We place the conversations
Synthetic callers with different voices, accents and moods run the identical set against every vendor, in parallel, recorded. Same tests, same order, same volume — the only reason your vendors are comparable at the end.
You get the graded report
A letter grade per vendor, the per-situation breakdown behind it, and the recording and transcript for every single conversation. Nothing is asserted that you can’t open and check yourself.
What you need before you sign up
This is the part other testing pitches leave out, so here it is first. We test real deployed products, which means the vendors have to let us in. Have these ready and a run starts the same day:
- A pilot or trial account with each vendor you want tested. Nearly every AI receptionist vendor offers one; ask for it while you’re negotiating anyway. Without it, all we can test is the vendor’s public demo line — which is what our free published grades already do, and which is not the same product you would deploy.
- One way to reach each vendor’s agent — the phone number, chat widget or email address it answers on. One per vendor, which is what we’ll test and re-test.
- Permission to test it. You confirm you’re authorized; we verify it again before every conversation we place. We do not call a line nobody agreed to have called.
- A seeded test record, if you want rescheduling or cancellation tested. To verify that an agent can move an appointment, an appointment has to exist. So your vendor (or you, in their system) creates one explicitly-marked test customer — fictional name, fictional date of birth, a number we control — and our callers reschedule and cancel against that record. No real customer, patient or account is ever touched.
What you get back
- A grade per vendor, on the same A–F scale as our public report cards — what the letters mean — with the pass rate and its range shown beside it, never a letter pinned tighter than the evidence supports.
- The scored line items, situation by situation: did it book the right slot, verify the right details, escalate the emergency, take a message that matched what the caller said.
- Every conversation, openable. Recording where there was one, transcript always, with the exact moments each score came from cited in it.
- The disclosures. Anything we couldn’t test, couldn’t observe, or that depended on how that vendor’s demo happens to be configured is labelled as exactly that and left out of the score — never silently converted into a failure.
Two things you won’t get anywhere else
A real head-to-head
Most “comparisons” are two separate reviews printed side by side. Ours compares vendors only on the situations they both ran, cell for cell — a vendor is never credited or penalised for a test the other one never took. Where the sets don’t match, we say they’re not comparable instead of pooling them into a number.
“Too close to call” is a result we’re willing to publish
When two vendors’ results are close enough that the difference could be the luck of which calls we happened to place, we tell you they’re indistinguishable rather than crowning the one that came out a hair ahead. That is not a hedge; it is the finding. It means the decision belongs on price, integration or support, and it means we’re not selling you a winner we can’t back. When a run is too small to separate them, we say so and tell you how many more conversations it would take.
How big a run needs to be
Everyone asks this second, right after what it costs, and the two are the same question: you pay per conversation, so the size of the run is the bill. A run is three numbers multiplied — vendors × situations × repeats:
- A quick first look: 2 vendors × 5 situations × 3 repeats = 30 conversations. Enough to see which one falls over; at that size grades usually come back as a range rather than a single letter.
- A typical shortlist run: 3 vendors × 8 situations × 5 repeats = 120 conversations. Where most buyers start, and normally enough to tell a good agent from a bad one.
- A decision you’re putting a year of budget behind: 3 vendors × 12 situations × 10 repeats = 360 conversations. The size at which two close vendors usually stop being “too close to call”.
Repeats are the part people want to cut, and they’re the part that decides whether the report can name a winner. One call tells you what happened once; five tell you whether it happens reliably. Eight situations run five times each is worth far more to you than forty situations run once.
Multiply your own three numbers and read the rate off the live plan cards on our pricing page — a fixed amount each month plus an amount per conversation, billed in arrears. (Working out what the vendors will charge you is a separate exercise: how much an AI receptionist costs.)
How long it takes
The conversations themselves are fast: hundreds of them run in parallel, so the calling is a matter of hours, not weeks. The honest variable is vendor access — how quickly each vendor grants your pilot and seeds a test record. After the calls, every failing result is re-reviewed before it reaches your report, and every judged pass or fail gets a second opinion from a judge at a different AI lab. We don’t hand you a number that hasn’t finished being checked.
Common questions
How is this different from the free grades on this site?
The published grades are earned on vendors’ public demo lines with generic scenarios — a fast, free way to shortlist. This is your shortlist, your situations, and the actual product each vendor would deploy for you, which a demo line can be either more polished or more limited than.
Do the vendors know which calls are yours?
They know a pilot is being exercised, because you asked them for one. They don’t get a schedule, a script or a heads-up, and our callers behave like the customers you actually get.
Can I use my own scenarios?
Yes — that’s the point of a private run. Write the situations your business actually gets, in plain language, and they’re run identically against every vendor on your list.
What if one vendor’s line goes down mid-run?
Conversations our platform broke are excluded and never billed. Conversations the vendor’s own system broke are a result — that is a real thing you would have lived with in production, and it counts.
Do you use real customer data?
Never. Callers are synthetic personas with clearly fictional details, and where a test needs an existing record, it runs against an explicitly-marked test record seeded for the purpose.
What does it cost?
A fixed amount per month plus an amount per conversation, billed in arrears, with the per-conversation rate stepping down as volume rises. For the quantity to multiply that rate by: a typical shortlist run is 3 vendors × 8 situations × 5 repeats = 120 conversations. The live plans are on the pricing page.
The independent guide to AI receptionists. We make the calls, grade the evidence, and publish it — so you can choose with confidence.