Your best jobs call after hours.
Can the AI handle them?
We grade AI answering services on the 2am burst-pipe call, not the sunny-day demo — dispatch, capture, and scheduling under pressure, all recorded.

Home-services report cards
Graded on the calls a trade actually gets — most of them after hours.
No grades published here yet. The test suite is built and waiting; we publish a grade only with the per-conversation evidence behind it.
Nothing graded here yet. When this page was built we had not published a grade for any vendor serving home-services businesses. That is a statement about our coverage, not a judgement about anyone: a public grade needs a vendor-published demo line and a round of calls clean enough to score, and we do not have both here yet. The suite below is built and waiting. The cards above are live — if one is showing, it published after this line was written, and the card is what counts.
What we test for home-services businesses
For plumbers, HVAC, electricians and restoration crews, the phone is the business — and most revenue calls come after hours. Our home-services rubric grades:
- After-hours emergencies — a burst pipe at 2am needs dispatch or a real escalation path, not voicemail.
- Job capture — address, problem, access details and callback number, recorded accurately.
- Scheduling windows — offering real arrival windows and confirming them back.
- Quote boundaries — collecting what an estimate needs without inventing prices.
- Caller patience — stressed callers interrupt and ramble; a good agent keeps up without repeating itself.
How to read the letter
A grade starts with how many calls went right, then moves down only for how the calls felt — long silences, repeated questions, talking over the caller.
The call count sits next to it
Twelve calls is not three hundred. Where the calls so far cannot narrow it to one letter, the card shows a range instead of pretending.
Some checks leave the count
What we could not test on a demo line, or what depends on how one office is set up, is removed and disclosed — never turned into a failure.
Nothing publishes on one opinion
Two separate checks, not one. Every failing result goes to a panel of three reviewers at three different AI labs, majority ruling. Every judged pass-or-fail also gets a second opinion from a judge at a different AI lab, and a disagreement goes to a person before anything can publish.
Whatever we have graded, here is your next step
Test your own shortlist
Point our synthetic callers at the vendors you are considering and get the same recorded, graded calls we run for this guide — your scenarios, your industry, priced per conversation.
Do it yourself in thirty minutes
The ten-question checklist is a condensed version of our rubric that works on any vendor’s demo line — including vendors we have never graded.
Read the grades we do have
Some vendors sell into several industries at once. The flagship guide lists every graded vendor we publish, across every industry.
Are you a vendor in this industry?
Verify with your work email and you can read the calls behind your results — every check, the reasoning, and the transcript — and dispute anything you think we got wrong. Disputes are free, always.
One line, so the numbers stay comparable. A provider account tests the demo line you publish — the same number your public grade is built from. That is deliberate: a retest only means something if it runs against the same endpoint and the same suite as the grade it updates. Want us to test a different number — an internal build, a staging agent, a configured customer account? That is a buyer account, and you can open one today; those results are yours and private, and they never become a public grade.
Common questions
Can an AI answering service really handle emergency calls?
Some can, most demos can’t. The test that matters: call the vendor’s line, describe water pouring through a ceiling, and see whether you reach dispatch or get offered Thursday at 3. That is literally one of our scored scenarios.
What does missing an after-hours call cost?
For most trades a single booked emergency pays for a month of answering — which is why we weight after-hours scenarios heavily in the grade.
Will it invent a price to get the job?
It should not. Collecting what an estimate needs — the problem, the access, the age of the equipment — is the passing behaviour; quoting a repair sight-unseen is not, and it is the fastest way to a cancelled job and a bad review.
Does it capture the address and access details accurately?
That is scored on its own, because a perfect booking at the wrong address is a wasted truck roll. We check the callback number and the access details come back the way the caller said them.
The independent guide to AI receptionists. We make the calls, grade the evidence, and publish it — so you can choose with confidence.