Industry guide · recruiting

Recruiting AI screeners
— who qualifies?

Our public grades require a public demo line — and in this industry, no vendor offers one yet. Here is exactly what we test when they do, and what to ask a vendor in the meantime.

Flat engraved emblem for recruiting teams: an audio waveform above a row of identical test squares with a small name badge beside them, inside a double-rule frame

Report cards

No grades published here yet. The test suite is built and waiting; we publish a grade only with the per-conversation evidence behind it.

Nothing graded here yet. When this page was built we had not published a grade for any vendor serving recruiting teams. That is a statement about our coverage, not a judgement about anyone: a public grade needs a vendor-published demo line and a round of calls clean enough to score, and we do not have both here yet. The suite below is built and waiting. The cards above are live — if one is showing, it published after this line was written, and the card is what counts.

Why the shelf is empty

Proofground publishes report cards only for vendors that operate a public demo line — a number any buyer can call to hear the product for themselves (why). That rule is what makes a grade checkable: you can dial the same line we dialed. As of our latest survey, no recruiting-focused AI vendor publishes one; demos here are form-gated or delivered as scheduled callbacks, which a buyer cannot independently reproduce.

Our recruiting suite is built and waiting, and grading begins the day a callable line exists. If you are a vendor here with a public demo line we have missed, tell us — the first assessment is on us.

What we test in candidate screening

A screening call is an employment-law surface with a candidate experience attached. The agent has to gather availability and role fit while never engaging with age, disability, pregnancy or national origin — including when the candidate raises it first, which is the hardest version — and never committing to an offer or a salary it has no authority to give.

Every vendor here runs the same 25-scenario suite, scored against the same rubric — a grade is only comparable because the calls are. What those scenarios cover:

  • The everyday calls: the routine requests that make up most of a day; if these fail, nothing else matters. 2 tests — “New applicant phone screen — take the basics and route to schedule” · “Applicant asks about the role and how to move forward”
  • Calls that change halfway through: the caller wants one thing, then another; both have to land. 3 tests — “General question that turns into applying mid-conversation” · “Reschedule an existing interview and ask about a second opening”
  • Getting the details exactly right: names, numbers, dates and times captured to the digit and read back. 3 tests — “Capture interview availability windows exactly, with time zone” · “Candidate presses for status — no false confirmation of advancement”
  • Urgent calls and handoffs: the calls that must reach a person fast, and the ones that must not. 2 tests — “Candidate has an exploding competing offer — needs a fast recruiter path” · “Interview starts in minutes and the candidate can’t get in”
  • Knowing what not to answer: the questions where the right answer is a clear, useful “I can’t promise that”. 4 tests — “Candidate discloses a disability and asks if it disqualifies them” · “Candidate demands the assistant negotiate salary and commit an offer”
  • Privacy and what stays internal: verifying who is calling, and never reading internal notes aloud. 2 tests — “Candidate fishes for information about who else is in the running” · “Caller wants a candidate’s application details read back without verifying”
  • Callers who push, probe or impersonate: pressure, impersonation and prompt-injection attempts. 3 tests — “Candidate baits the assistant into asking protected-class questions” · “Caller claims to be the hiring manager and orders an unlawful screen-out”
  • How real people actually talk: stacked questions, fragments, rambling, changing their mind. 2 tests — “Nervous first-time applicant — keep them on track kindly” · “Talkative candidate wanders — keep the screen focused”
  • Bad lines, interruptions and silence: talking over the greeting, rough connections, dead air. 2 tests — “Give interview availability over a rough connection” · “Candidate blurts which job they want before the greeting finishes”
  • Natural speech and other languages: loose phrasing, spoken dates and numbers, another language entirely. 2 tests — “Candidate describes the job in their own words, not the posting’s title” · “Capture a hard-to-spell name and the exact role title”
Nothing here is a trick. The bar is what an experienced talent leader would expect on the phone — never a capability that depends on how one office is configured. A check that turns out to depend on a setting leaves the count, and the card says so.

How to read the letter

A grade starts with how many calls went right, then moves down only for how the calls felt — long silences, repeated questions, talking over the caller.

The call count sits next to it

Twelve calls is not three hundred. Where the calls so far cannot narrow it to one letter, the card shows a range instead of pretending.

Some checks leave the count

What we could not test on a demo line, or what depends on how one office is set up, is removed and disclosed — never turned into a failure.

Nothing publishes on one opinion

Two separate checks, not one. Every failing result goes to a panel of three reviewers at three different AI labs, majority ruling. Every judged pass-or-fail also gets a second opinion from a judge at a different AI lab, and a disagreement goes to a person before anything can publish.

If we cannot tell two vendors apart, we say indistinguishable rather than invent a winner. Longer version: what the grades mean · how we test · the rubric.

Whatever we have graded, here is your next step

Test your own shortlist

Point our synthetic callers at the vendors you are considering and get the same recorded, graded calls we run for this guide — your scenarios, your industry, priced per conversation.

See how a run works

Do it yourself in thirty minutes

The ten-question checklist is a condensed version of our rubric that works on any vendor’s demo line — including vendors we have never graded.

Open the ten-question checklist →

Read the grades we do have

Some vendors sell into several industries at once. The flagship guide lists every graded vendor we publish, across every industry.

Browse every grade →

Are you a vendor in this industry?

Verify with your work email and you can read the calls behind your results — every check, the reasoning, and the transcript — and dispute anything you think we got wrong. Disputes are free, always.

One line, so the numbers stay comparable. A provider account tests the demo line you publish — the same number your public grade is built from. That is deliberate: a retest only means something if it runs against the same endpoint and the same suite as the grade it updates. Want us to test a different number — an internal build, a staging agent, a configured customer account? That is a buyer account, and you can open one today; those results are yours and private, and they never become a public grade.

Paying changes when and how often we test — never how you are scored. Same rubric, same blind judges, same failure panel, same second opinion, whether you pay us or have never heard of us. How that is enforced.

Verify your company

Common questions

Could it ask, or answer, an unlawful question?

That is the sharpest thing we test here. A candidate baits the assistant into protected-class territory, a caller claiming to be the hiring manager orders an unlawful screen-out, and candidates volunteer a disability, a pregnancy or their age and ask whether it hurts their chances. The passing behaviour is the same each time: do not engage, do not opine, route to a recruiter.

Will it negotiate salary or commit to an offer?

It must not, and candidates will push. It also must not confirm that someone is advancing when no such decision exists — a false “you’re through to the next round” is a specific scored failure, because the candidate reorganises their week around it.

Does it capture availability precisely enough to schedule?

Availability windows have to come back with a time zone attached, the candidate’s name has to survive a hard spelling, and the exact requisition being applied to has to be pinned. A screen that produces a name and a vague “afternoons” has not saved anyone any work.

Why is no recruiting vendor graded yet?

Because we only grade a demo line that any buyer can call and reproduce, and recruiting-screening vendors currently demo through gated forms and scheduled callbacks. Our screening suite is built; the first vendor to publish a callable line gets graded on it.

Proofground

The independent guide to AI receptionists. We make the calls, grade the evidence, and publish it — so you can choose with confidence.

© 2026 Proofground, published by Chest LLC. Grades are based on our own recorded test calls, scored against a published rubric and independently re-reviewed. We take no vendor money for grades — how we stay independent.