The idea behind Proofground

Mystery shoppers,
for the AI age

Businesses have mystery-shopped their stores for decades. We do it to AI receptionists — recorded calls, blind grading, published evidence.

Flat engraved emblem: an audio waveform above a row of identical test squares with a small magnifying glass beside them, inside a double-rule frame

What mystery shopping is, and why it fits AI

Mystery shopping is a plain old idea: you pay someone to be a customer. They walk in, order the thing, ask the awkward question, and write down what actually happened — not what the training manual says happens. Businesses have done it for decades, because the only honest way to know what a customer experiences is to go and experience it.

It transfers to an AI phone agent almost perfectly, and in one way it fits better than it ever fitted a shop floor: a visit is expensive and every shopper is a different person having a different day, while an AI receptionist is an encounter you can sample over and over, identically, with callers you define down to the accent and the mood.

What changes when the thing you are shopping is software

A human shopper comes back with a story: I called on a Tuesday, it put me on hold, I gave up. True, and nearly useless for choosing a vendor, because you cannot know whether Tuesday was typical. Place that call again and again and change one thing — background noise, a stronger accent, a caller who interrupts, the same request at two in the morning — and you can see what that change did. What comes back is not an anecdote but a measurement: how often the booking landed, how often the caller had to repeat themselves, what went wrong on the calls that did. That is all we do, over and over, on the record.

Why the “mystery” part matters: a vendor never knows which calls are ours, so what we measure is the product every caller gets — not a version warmed up for a tester.

What a test call is made of

Two things, deliberately kept apart.

The scenario — what the caller wants

Book a first appointment. Move Thursday to next week. Cancel. Report something urgent. A scenario is a goal and nothing else — never a word about the voice on the line.

The persona — who is calling

An accent. A television in the background. Someone who mumbles, talks across the agent, or rings already annoyed. A persona is a way of speaking and nothing else — never a reason for the call.

Why they are never mixed: run one rescheduling request with a clear caller and again on a noisy line, and the difference belongs to the noise. That is what lets a card say this agent handles rescheduling but struggles on noisy calls rather than that call went badly. Fold the noise into the scenario and a failure has two possible causes with no way to tell them apart.

Our caller follows the agent’s lead

A tester who opens with “Hi, I’m Jane Doe, date of birth the fourth of March, my insurance is…” makes almost any agent look capable. Real callers do not do that, so ours does not either: it says why it is ringing, answers what it is asked, and volunteers nothing. If the agent never asks for a callback number it never gets one, and that counts as its miss. What gets graded is the agent’s own flow, not a script that did half the job for it.

How a call becomes a grade

Each call is scored against a written rubric — published, so you can see what was asked of a product before you read what it earned — by judges who are not told whose product it is. Then two separate checks run.

Every failing result is re-examined before publication by a panel of three reviewers, each running at a different AI lab, asking one question: was this genuinely the product’s failure — or our caller’s, or a limit of the demo line we called? The majority decides, and a split panel leaves the failure standing.

Every judged pass-or-fail is then judged a second time by a model at a different AI lab. Clear agreement settles it. A disagreement — or a score sitting too close to the pass line for agreement to mean much — goes to a person, and nothing resting on that result publishes until they have ruled.

“We could not measure that” is not a failure

The line we hold hardest is between a product that failed and something we could not observe. A demo line with no route to a real calendar has not failed at booking; behaviour that depends on one office’s settings is neither the product’s fault nor its credit; a call spoiled by our own equipment says nothing about the vendor. Those results leave the count and are named on the card, never counted as failures.

Three reasons people mystery-shop an AI

You are choosing between vendors

Start with the published grades; if the vendors on your list are not among them, the same calls can be run against your shortlist.

You already run an AI on your own phone line

The underestimated case: the agent on your line changes without asking you — a model update, a new prompt, a setting someone adjusted — and the first sign of trouble is usually a customer mentioning it. Shopping on a schedule catches that early, because identical tests run before and after.

That makes you a buyer here, not a vendor — the vendor door checks your email domain against the vendor’s website and would turn you away. See what a run involves, then open a buyer account: same callers, same rubric, your number, results private to you.

You sell the AI

Continuous shopping catches a regression before your customers report it — an update that starts talking over callers, a prompt change that breaks rescheduling. See how the vendor side works. Paying changes when and how often we test — never how you are scored.

Shop a demo line yourself, this afternoon

Take a vendor’s published demo number, set aside twenty minutes, and listen for the following — writing each down, so you end up with something comparable rather than an impression.

  1. Count the silence before it speaks. A second or two is normal. Longer, and real callers start hanging up.
  2. Interrupt it mid-sentence. A good agent stops and listens; a weak one finishes its script first.
  3. Correct yourself on purpose. “Tuesday — sorry, Wednesday.” Listen for which day comes back; an agent that hears only your first answer fails silently here.
  4. Give it a name and a number to read back. A surname it cannot have guessed, said at normal speed. Mangled names and numbers make work.
  5. Call from somewhere noisy. A corridor, a car, a television. Does it ask you to repeat, as a person would — or guess?
  6. Ask to change a booking, not just make one. Booking is the rehearsed path. Rescheduling makes it find you first, and thin products come apart there.
  7. Ask something it should not answer. A price, a clinical question. A clean handoff is right; a confidently invented answer is the costliest failure here.
  8. Say something urgent in a real caller’s words. Not “this is an emergency” but the real sentence — water through the ceiling. Anything short of escalation is disqualifying.
  9. Then call again tomorrow. One call is a story. Two that differ tell you more than either, and that gap is what you are buying.
What to ask the vendor before you sign is a different list: the ten questions. And if a vendor already has a report card, most of this has been run for you.

Common questions

What is mystery shopping for AI?

The same discipline as a classic mystery shopping programme — a shopper who behaves like a customer, standard scenarios, a scored report — except the shopper is synthetic and the subject is an AI phone agent.

How is it different from hiring a mystery shopping company?

A person makes a handful of calls, each slightly different, because people are. A synthetic caller repeats one call with a single thing changed, and runs at several vendors with identical scenarios — which is the only reason two grades are comparable.

How is this different from a vendor testing itself?

A vendor testing itself grades its own homework. We choose the scenarios, our judges are blind to the vendor, and results publish whether the vendor likes them or not — after the two independent checks on how we test.

Can I mystery-shop the AI that answers my own phone line?

Yes, and that is a buyer account, not a vendor one. The same callers and rubric point at your number on a schedule, so a change your vendor ships shows up between two runs of identical tests.

Proofground

The independent guide to AI receptionists. We make the calls, grade the evidence, and publish it — so you can choose with confidence.

© 2026 Proofground, published by Chest LLC. Grades are based on our own recorded test calls, scored against a published rubric and independently re-reviewed. We take no vendor money for grades — how we stay independent.