Industry guide · freight

Freight AI dispatch agents,
put to the test

The same discipline as every Proofground guide: recorded calls, blind judges, published evidence — on the scenarios this trade lives on.

Flat engraved emblem for freight brokers: an audio waveform above a row of identical test squares with a small trailer beside them, inside a double-rule frame

Report cards

Graded on the calls freight brokers actually get — recorded, judged against a written rubric, independently re-reviewed.

Grades could not be loaded just now. Please try again in a moment.

Graded so far: CloneOps.ai. Open a card for the letter, the calls behind it, and the date they were made — every grade is a dated snapshot, not a standing verdict, and the cards above are the current list.

What we test in freight dispatch

Dispatch is reference numbers and authority. A load number, a BOL, a PO and a PRO are four different things that must not merge into one field, an appointment time belongs to the shipper’s timezone and nobody else’s, and rate, detention and lumper pay are commitments an agent usually has no authority to make. On top of that sits the call where a driver is in trouble.

Every vendor here runs the same 26-scenario suite, scored against the same rubric — a grade is only comparable because the calls are. What those scenarios cover:

  • The everyday calls: the routine requests that make up most of a day; if these fail, nothing else matters. 2 tests — “Owner-operator books a posted load — clean end-to-end dispatch” · “Driver makes a routine check call on a load already in progress”
  • Calls that change halfway through: the caller wants one thing, then another; both have to land. 3 tests — “Booking a new load pivots into a detention-pay problem on the current one” · “Rate, pickup appointment, and status of another load — all in one breath”
  • Getting the details exactly right: names, numbers, dates and times captured to the digit and read back. 5 tests — “BOL number and PO number kept separate, not merged into one field” · “Delivery appointment confirmation number captured and read back”
  • Urgent calls and handoffs: the calls that must reach a person fast, and the ones that must not. 3 tests — “Driver in an accident with a possible injury — safety first, then the load” · “Truck broken down on the shoulder with a loaded trailer — route to real help now”
  • Knowing what not to answer: the questions where the right answer is a clear, useful “I can’t promise that”. 3 tests — “Pushed to lock a rate the agent has no authority to commit” · “Won’t invent a delivery appointment slot it doesn’t control”
  • Privacy and what stays internal: verifying who is calling, and never reading internal notes aloud. 2 tests — “Internal TMS codes and carrier-rating notes read aloud to the caller” · “Fishing for another carrier’s rate, the broker’s margin, and the shipper’s identity”
  • Callers who push, probe or impersonate: pressure, impersonation and prompt-injection attempts. 2 tests — “Caller tries to access and reroute a load that isn’t theirs” · “Caller tries to jailbreak the dispatch agent out of its role”
  • How real people actually talk: stacked questions, fragments, rambling, changing their mind. 2 tests — “’Yeah’ to a two-in-one question — does the agent clarify or guess” · “Load and truck details given in scattered fragments — does the agent re-ask what it has”
  • Bad lines, interruptions and silence: talking over the greeting, rough connections, dead air. 2 tests — “Dead air while the agent pulls up the load — does it manage the silence” · “Driver rattles off load info before the greeting finishes — is any of it caught”
  • Natural speech and other languages: loose phrasing, spoken dates and numbers, another language entirely. 2 tests — “Driver’s natural time-and-window phrasing resolved to a concrete appointment” · “Whole dispatch call expected in the caller’s preferred language — does the agent stay and confirm”
Nothing here is a trick. The bar is what an experienced dispatcher would expect on the phone — never a capability that depends on how one office is configured. A check that turns out to depend on a setting leaves the count, and the card says so.

How to read the letter

A grade starts with how many calls went right, then moves down only for how the calls felt — long silences, repeated questions, talking over the caller.

The call count sits next to it

Twelve calls is not three hundred. Where the calls so far cannot narrow it to one letter, the card shows a range instead of pretending.

Some checks leave the count

What we could not test on a demo line, or what depends on how one office is set up, is removed and disclosed — never turned into a failure.

Nothing publishes on one opinion

Two separate checks, not one. Every failing result goes to a panel of three reviewers at three different AI labs, majority ruling. Every judged pass-or-fail also gets a second opinion from a judge at a different AI lab, and a disagreement goes to a person before anything can publish.

If we cannot tell two vendors apart, we say indistinguishable rather than invent a winner. Longer version: what the grades mean · how we test · the rubric.

Whatever we have graded, here is your next step

Test your own shortlist

Point our synthetic callers at the vendors you are considering and get the same recorded, graded calls we run for this guide — your scenarios, your industry, priced per conversation.

See how a run works

Do it yourself in thirty minutes

The ten-question checklist is a condensed version of our rubric that works on any vendor’s demo line — including vendors we have never graded.

Open the ten-question checklist →

Read the grades we do have

Some vendors sell into several industries at once. The flagship guide lists every graded vendor we publish, across every industry.

Browse every grade →

Are you a vendor in this industry?

Verify with your work email and you can read the calls behind your results — every check, the reasoning, and the transcript — and dispute anything you think we got wrong. Disputes are free, always.

One line, so the numbers stay comparable. A provider account tests the demo line you publish — the same number your public grade is built from. That is deliberate: a retest only means something if it runs against the same endpoint and the same suite as the grade it updates. Want us to test a different number — an internal build, a staging agent, a configured customer account? That is a buyer account, and you can open one today; those results are yours and private, and they never become a public grade.

Paying changes when and how often we test — never how you are scored. Same rubric, same blind judges, same failure panel, same second opinion, whether you pay us or have never heard of us. How that is enforced.

Verify your company

Common questions

Will it commit to a rate it has no authority to give?

It should not, and carriers will push. We test a caller locking for a rate, another pressing for guaranteed detention and lumper pay, and a third pushing for a delivery appointment slot the broker does not control. Holding the line politely and routing to a person is the passing behaviour.

Does it keep BOL, PO and PRO numbers straight?

That is the biggest group in this suite. Numbers have to stay in their own fields, come back without a transposed digit on read-back, and a pickup appointment has to be pinned to the shipper’s local timezone rather than whatever the agent assumed.

What happens on an accident or a hazmat leak?

Safety comes before the load, and we score whether the agent behaves that way. A driver in an accident with a possible injury, a truck broken down on the shoulder loaded, and a placarded hazmat load leaking each require real emergency routing — not a ticket in a queue.

Can it handle a routine check call without a human?

That is the baseline, along with an owner-operator booking a posted load end to end. We also measure what the agent does with its own dead air while it looks a load up — silence on a dispatch line is how a driver decides to call a person instead.

Proofground

The independent guide to AI receptionists. We make the calls, grade the evidence, and publish it — so you can choose with confidence.

© 2026 Proofground, published by Chest LLC. Grades are based on our own recorded test calls, scored against a published rubric and independently re-reviewed. We take no vendor money for grades — how we stay independent.