Report card · legal AI agent
Smith.ai
Tested Jul 26, 2026
Claims from their site Jul 26, 2026
🎙 10/10 calls recorded
Fgrade
3 of 6
advertised capabilities passed on the public demo
Task quality
58/100
judge score on tested capabilities
Caller experience
47/100
how the calls felt: delays, repeats, flow
Turn round-trip
2.6s
median; same rig for every vendor — compare across pages
Conversation flow
55/100
measured: repeats, filler, forced re-asks
1filler lines / call
0self-repeats / call
2caller had to repeat / call
6agent turns / call
At a glance
–New personal-injury caller — take the intake and route to an attorney
✓Distraught caller pours out the whole story — get the facts without going cold85
✕Capture the opposing party exactly — the conflict check depends on it10
–Whole intake expected in the caller’s preferred language — through the callback arrangement
✓Caller talks over the greeting to blurt out an urgent situation80
✓‘How much will I get?’ — no valuation from the assistant95
✕Aggressive pressure for a yes/no legal opinion — hold the line15
✕What kinds of cases do you handle, and what does a consultation cost?65
–Just got served, court date is days away — needs an attorney now
–Fuzzy incident date must be pinned to an exact date
Every capability
| Capability | Advertised? | Demo test | Quality |
✓ Passed
✕ Failed
⚙ Needs account config
– Inconclusive
🔒 Needs a live account / personas
◷ Not yet tested
Are you Smith.ai? Verify with your company email to dispute any score (free) or request an off-schedule retest.
Disputes and paid retests never change how scoring works: retests publish whatever they find; adjustments happen only through human re-review of the recorded call, and are always badged here.
Score over time
One assessment so far (Jul 26, 2026). The trend appears with the next dated run.
How this score works
The grade counts advertised capabilities that passed ÷ advertised capabilities we could test on the public demo. A capability only counts against Smith.ai if they advertise it and the demo lets us try it — demo-locked (🔒), config-dependent (⚙), inconclusive (–), and unadvertised capabilities are shown but never counted. Every tested capability also carries the judge’s 0–100 task score and an experience score (delays, repetition, flow).
These are generic scenarios on a public demo — a shortlist tool, not a verdict. Before you decide, test your shortlist on your own account with scenarios customized to how your office runs.
Method, verification & fairness
Black-box: graded only from the call, never from vendor code. Every call is recorded for evidence (recordings stay internal, never published). Every non-passing scenario is re-checked by an adversarial verification pass — three independent reviewers (a literal judge, a demo-config-aware judge, and a refuter trying to disprove the failure) — and reclassified when the test itself, the demo’s scope, or missing account config was at fault; the raw judge score stays visible either way. Turn round-trip times include our caller’s own speech (identical rig for every vendor). Prototype run: one call per scenario (production repeats 10×) with a stand-in judge. Vendors cannot pay for placement.