Hotel AI reservation agents,
put to the test
The same discipline as every Proofground guide: recorded calls, blind judges, published evidence — on the scenarios this trade lives on.

Report cards
Graded on the calls hotels actually get — recorded, judged against a written rubric, independently re-reviewed.
No grades published here yet. The test suite is built and waiting; we publish a grade only with the per-conversation evidence behind it.
Graded so far: roommaster (InnQuest Software). Open a card for the letter, the calls behind it, and the date they were made — every grade is a dated snapshot, not a standing verdict, and the cards above are the current list.
What we test in hotel reservations
A reservation line is a promise machine: a rate, a room type, a number of nights, an arrival time. Every one of those can be wrong in a way the guest only discovers at the desk, and a false yes on an accessible room or a pet policy costs far more than a lost booking. So we grade arithmetic and honesty as hard as we grade booking.
Every vendor here runs the same 27-scenario suite, scored against the same rubric — a grade is only comparable because the calls are. What those scenarios cover:
- The everyday calls: the routine requests that make up most of a day; if these fail, nothing else matters. 2 tests — “Front-desk info: check-in time, parking, and amenities” · “Straightforward availability check and nightly rate quote”
- Calls that change halfway through: the caller wants one thing, then another; both have to land. 4 tests — “Mid-booking, the caller pivots to changing an existing reservation” · “Calls to cancel, then decides to move the dates instead”
- Getting the details exactly right: names, numbers, dates and times captured to the digit and read back. 4 tests — “’Is it actually booked?’ — a truthful yes, with a real confirmation” · “’Next Friday’ — which Friday? Resolve, don’t assume”
- Urgent calls and handoffs: the calls that must reach a person fast, and the ones that must not. 2 tests — “Billing dispute the agent can’t resolve — get it to a manager” · “Sold out for the dates — escalate or waitlist, don’t invent a room”
- Knowing what not to answer: the questions where the right answer is a clear, useful “I can’t promise that”. 4 tests — “ADA accessible room needed — honest availability, no false yes” · “Wants a guaranteed suite upgrade the agent can’t promise”
- Privacy and what stays internal: verifying who is calling, and never reading internal notes aloud. 2 tests — “Asks for details of someone else’s reservation” · “Probing for internal rate codes and system details”
- Callers who push, probe or impersonate: pressure, impersonation and prompt-injection attempts. 2 tests — “Pressed on whether it’s a real person — must not pretend” · “Jailbreak attempt to force a free upgrade and waive fees”
- How real people actually talk: stacked questions, fragments, rambling, changing their mind. 3 tests — “A single ‘yes’ to a two-part question — which part?” · “Still angry about the last stay but wants to book again”
- Bad lines, interruptions and silence: talking over the greeting, rough connections, dead air. 2 tests — “Confirmation-name on the reservation garbled and must be gotten right” · “Caller starts the request before the greeting finishes”
- Natural speech and other languages: loose phrasing, spoken dates and numbers, another language entirely. 2 tests — “Asks to be helped in their language or by someone who speaks it” · “Whole reservation expected in the caller’s preferred language — including the rate and confirmation”
How to read the letter
A grade starts with how many calls went right, then moves down only for how the calls felt — long silences, repeated questions, talking over the caller.
The call count sits next to it
Twelve calls is not three hundred. Where the calls so far cannot narrow it to one letter, the card shows a range instead of pretending.
Some checks leave the count
What we could not test on a demo line, or what depends on how one office is set up, is removed and disclosed — never turned into a failure.
Nothing publishes on one opinion
Two separate checks, not one. Every failing result goes to a panel of three reviewers at three different AI labs, majority ruling. Every judged pass-or-fail also gets a second opinion from a judge at a different AI lab, and a disagreement goes to a person before anything can publish.
Whatever we have graded, here is your next step
Test your own shortlist
Point our synthetic callers at the vendors you are considering and get the same recorded, graded calls we run for this guide — your scenarios, your industry, priced per conversation.
Do it yourself in thirty minutes
The ten-question checklist is a condensed version of our rubric that works on any vendor’s demo line — including vendors we have never graded.
Read the grades we do have
Some vendors sell into several industries at once. The flagship guide lists every graded vendor we publish, across every industry.
Are you a vendor in this industry?
Verify with your work email and you can read the calls behind your results — every check, the reasoning, and the transcript — and dispute anything you think we got wrong. Disputes are free, always.
One line, so the numbers stay comparable. A provider account tests the demo line you publish — the same number your public grade is built from. That is deliberate: a retest only means something if it runs against the same endpoint and the same suite as the grade it updates. Want us to test a different number — an internal build, a staging agent, a configured customer account? That is a buyer account, and you can open one today; those results are yours and private, and they never become a public grade.
Common questions
Will it quote the right rate and the right number of nights?
We check the arithmetic directly: three nights versus the checkout day is a classic off-by-one, and “about a hundred and something” is not a rate quote. The agent has to give the exact nightly rate and total, and resolve “next Friday” to a real date rather than assuming which one the guest meant.
How does it handle accessible rooms, pets and early check-in?
These are the questions where a confident wrong answer does the damage. An ADA accessible room needs honest availability — not a reflexive yes — and a guaranteed suite upgrade or early check-in the agent cannot commit to has to be described as what it is: a request, not a promise.
Does it actually complete a booking, or just say it did?
We ask the blunt question every guest eventually asks: “is it actually booked?” The agent must give a truthful yes with a real confirmation, or a truthful no. Claiming a booking that does not exist is one of the fastest ways to fail a call on this suite.
What about an angry guest, or dates that are sold out?
A guest still annoyed about their last stay but wanting to book again is a scored scenario — staying warm and still completing the booking. Sold-out dates must produce a waitlist or an escalation, never an invented room, and a billing dispute has to reach a manager rather than being absorbed by the agent.
The independent guide to AI receptionists. We make the calls, grade the evidence, and publish it — so you can choose with confidence.