Guide · 10 min read · August 25, 2026
How to measure whether an AI receptionist is working: transcripts, bookings, escalations
The handful of numbers and the weekly transcript review that tell you whether an AI receptionist is helping callers or quietly losing them.
Measure an AI receptionist on outcomes, not on how many calls it answered. Track the share of calls that ended in a completed booking or a delivered message, the share escalated to a person and why, the share where the caller hung up before anything was completed, and the accuracy of the bookings it made. Pair those numbers with a weekly read of a sample of transcripts, because the numbers tell you something is wrong and the transcripts tell you what.
Start from the outcome each call was supposed to have
The easiest number to report about an AI receptionist is the number of calls it answered, and it is nearly useless. A receptionist that answers every call and completes none of them is worse than a busy signal, because callers believe they have been dealt with.
The measurement that matters starts from a simple observation: every call had an intended outcome. The caller wanted an appointment, an answer, a message delivered or a person. The receptionist either produced that outcome, produced a different acceptable one (a message instead of a transfer at 9 p.m.), or produced nothing. Everything below is a way of counting those three buckets and finding out why the third one is not empty.
The five numbers worth tracking weekly
Keep the dashboard small. Five figures, reviewed weekly, will show most problems within a fortnight of their appearing.
- Completion rate: the share of calls that ended with a booking made, a question answered from the script, or a message captured and delivered. This is the headline number.
- Escalation rate and reasons: the share of calls handed to a person, broken down by trigger — caller asked, urgency words, misunderstanding, out-of-scope. A high “misunderstanding” share points at the script; a high “out-of-scope” share points at jobs the receptionist should be given.
- Abandon-before-outcome rate: calls where the caller hung up after the greeting but before any outcome. Watch where in the call they left; drop-off during the greeting means it is too long, drop-off during questions means there are too many.
- Booking accuracy: of the appointments the receptionist made, how many needed correction by staff afterwards — wrong duration, wrong provider, wrong day. Sample this from the calendar, not from the receptionist’s own log.
- Time to human on escalations: from the moment the receptionist decided to transfer to the moment a person answered or a message was delivered. Long times here mean the destination or fallback chain is wrong, not the receptionist.
Read the transcripts — a sample, every week
Numbers show that something is off. Transcripts show what. Set aside thirty minutes a week for the first two months, and fifteen minutes a week after that, to read a sample of transcripts chosen deliberately: every call that ended in abandonment, a handful of escalations of each type, and a few random completed calls to check the receptionist is not succeeding by accident.
Read for specific things. Did the receptionist confirm its interpretation of the request before acting? Did it read the booking back? Did it ask for information the job did not need? Did it say anything outside its script — a price, a promise, an opinion? Did the caller repeat themselves, and if so, which word or phrase was the receptionist missing?
Keep a short running list of findings and the change each one implies. Most findings are script edits: a synonym to add, a question to remove, a service name the receptionist keeps mishearing. A few are routing changes. Very few are reasons to turn the receptionist off.
Check bookings against the calendar, not against the log
The receptionist will report the bookings it believes it made. The truth lives in the calendar and at the front desk. Once a week, pull the appointments the receptionist created and ask the people who manage the schedule how many they had to touch. A booking that was made at the right time but for the wrong duration, or with the wrong provider, or for a new patient in a returning-patient slot, still counts as an error even though the receptionist logged it as a success.
Also look for the reverse: bookings that did not happen. A caller who wanted an appointment and ended up leaving a message because the receptionist could not find a slot, or misheard the service, is a missed booking. Those show up in the escalation and message transcripts, and they are the most valuable thing to fix, because each one was a customer ready to commit.
Look at escalations from the staff side
An escalation is only successful if the person who received it thought it was appropriate and had what they needed. Ask the staff who take escalated calls two questions each week: were there calls you received that the receptionist should have handled, and were there calls it handled that should have come to you?
The first answer tends to shrink the escalation list — staff notice that “can I get directions” or “are you open Saturday” keeps arriving as a transfer. The second answer is more important and rarer; it usually surfaces a category of caller — a distressed patient, a commercial account with a complaint — that the receptionist is trying to serve when it should be stepping aside. Adjust the triggers, and tell the staff what changed.
Also ask whether the context that arrived with the transfer was useful. If staff are re-asking the caller’s name and reason every time, the whisper announcement or message summary is not doing its job.
What good looks like, and when to worry
There is no universal benchmark for these numbers, because a clinic taking bookings all day and a law firm taking messages after hours are measuring different jobs. What you can do is set a baseline in the first month and watch the direction. Completion rising, abandonment falling, misunderstanding escalations falling, and booking corrections falling is the pattern of a receptionist that is being tuned well.
Worry when the numbers are flat and the transcripts keep showing the same findings, because that means the changes are not being made. Worry more when completion is high but staff report corrections and re-asking, because that means the receptionist is claiming success it is not delivering. And worry immediately about any transcript where an urgent caller was kept in the script; that is a rule to fix the same day.
How long should we wait before judging whether an AI receptionist is working?
Give it four to six weeks with weekly reviews, because the first two weeks are spent finding script gaps that only real callers reveal. Judge it on the trend across that period rather than on the first week’s numbers. A receptionist whose completion rate and booking accuracy are improving each week is working; one that is flat after six weeks of adjustments probably has a scope problem — it has been given jobs that need a person.
Which single number should the owner look at if there is only time for one?
Abandon-before-outcome: the share of callers who hung up without a booking, an answer, a message or a transfer. It is the closest thing to a count of customers lost. Completion rate can look healthy while callers give up, and escalation rate can be high for perfectly good reasons. If abandonment is rising, read the transcripts of those calls that week; the cause is almost always visible within the first few.
Want to see your own call flow?
Three minutes of questions, a recommended setup.
Build My Phone System