What an AI receptionist can and cannot do on a call
An honest map of where a voice agent is genuinely better than a person, where it is worse, and where the handover has to sit.

The first question anyone asks about an AI receptionist is whether it sounds human. It's the wrong question, and it's the one every vendor answers, which tells you something.
The right question is narrower and much more useful. Which of your calls can this thing actually finish, and what happens on the ones it can't?
We've built and tuned enough of these now to have a clear answer, including the parts that don't flatter the technology. Here's the honest map.
#Where a voice agent is genuinely better than a person
Not "as good as". Better, and not by a small margin.
Pickup speed. It answers on the first ring, every time, including at 2am and including on the fourth simultaneous call. A human front desk cannot do this and it's unreasonable to expect them to. Given how much the first five minutes decide about who wins the job, this one advantage carries most of the business case on its own.
Concurrency. Ten people calling at once is a disaster for a two person office and a non-event for an agent. This matters more than owners expect, because enquiries arrive in clumps. A van breaks down, a storm passes through, an ad goes live, and suddenly the phone rings six times in four minutes.
Consistency of the record. A person writes a good note when they have time. An agent writes the same structured record on call four hundred as on call one, with the transcript attached. Six months later, when you want to know why bookings from one postcode never convert, that difference is the whole difference.
Never getting bored of the same question. Roughly half the calls into a clinic or a trades business are the same eight questions. Opening hours, parking, do you cover my area, what does a callout cost, can I move my appointment. A person answering those for the ninth time that morning is a person not doing anything more valuable.
Patience with slow callers. This one surprises people. An agent doesn't sigh, doesn't rush an elderly caller, and doesn't get shorter at 4:55pm on a Friday.
Working in more than one language without a rota. Ours run in six. For a clinic with a mixed patient list, that removes a staffing constraint that was previously solved by hoping the right person was on shift.
#Where it's worse, and will stay worse
Judgement calls with real stakes. A caller who wants a discount, a customer threatening to leave, a dispute about work already done. These need someone who can decide something and mean it. An agent that improvises here will either give away money or make an enemy.
Emotionally loaded conversations. A distressed patient, a bereavement, a caller who is frightened about a bill. The technology can produce sympathetic words. That isn't the same as being the right thing to put in front of someone at the worst moment of their week, and we advise clients to route these to a person immediately.
Anything that depends on undocumented context. "Put me through to the lad who came out last time." An agent knows what's in your systems. If half your operating knowledge lives in one person's head, the agent will hit that wall and so will the caller.
Accents, noise and bad lines. Recognition has improved a great deal and it is still not perfect. A caller on a building site with wind and machinery is a harder problem than a quiet kitchen, and the failure mode has to be a clean handover rather than three rounds of "sorry, could you repeat that".
Selling something the caller didn't ring about. A good salesperson hears an opening and takes it. An agent works to a script and a set of rules, so it will book the job the caller asked for and miss the larger one sitting behind it. On high value work, that gap is a real cost and it belongs on the honest side of the ledger.

#The escalation rule matters more than the script
Most people building one of these spend their time on the script. That part is easy and it changes weekly anyway. Escalation is what decides whether this is safe to put on your main number.
Write it before anything else, and write it as a list of triggers rather than a vague principle. Ours usually include:
- The caller asks for a person, in any wording, at any point. No second attempt to handle it.
- Any mention of a complaint, a refund, legal action, or an injury.
- Two failed attempts at understanding the same thing.
- Anything outside the defined scope, including any question about money that isn't a published price.
- Detected distress in how the caller is speaking, not just in the words.
Then decide what escalation actually does at each hour of the day. A warm transfer at 11am and a booked callback at 11pm are both fine. Silence is not, and neither is a promise that nobody has been told about.
Worth saying plainly: an escalation is not a failure. We track escalation rate as a health metric rather than an error rate, and a figure of zero would worry us more than a figure of fifteen percent. Zero usually means the agent is bulldozing through calls it should have handed over, and you find out about it in a review six weeks later.
This boundary is a feature you can sell, not an admission. It's precisely what makes an agent safe to deploy in a clinic or a law firm, and buyers in those sectors ask about it before anything else.
#Latency is what gives it away, not vocabulary
If you take one technical point from this, take this one.
People do not decide they're talking to a machine because of word choice. They decide because of the gap. Human conversational turn-taking runs on a rhythm of roughly two hundred milliseconds. Push the response gap past about a second and something feels wrong, even to a caller who couldn't tell you what changed. Push it past two and they start talking over the agent, which breaks the turn structure entirely and makes everything after it worse.
This is why we tune for response time before we tune for phrasing. It's also why a demo that sounds impressive in a quiet room can fall apart on a real phone line, where network conditions add their own delay. Ask any vendor for a recording of a live call on a mobile, outdoors. The difference from the showreel is usually instructive.
#How to start without betting the phone number on it
The failure pattern we see most often is scope. Someone buys an agent, points it at the main line, gives it everything, and it does eleven things at a mediocre standard.
The alternative takes longer to feel impressive and works far better.
Pick one call type. High volume, low judgement. Appointment booking and rescheduling is the usual answer. "What are your opening hours and do you cover my postcode" is a good one too, and it's often a third of the calls.
Point it at overflow first, not the main line. Calls that would otherwise ring out. The comparison is then against a voicemail rather than against your best receptionist, which is both fairer and much less risky.
Read the transcripts every week for the first month. Not a sample. All of them. This is the step everyone skips and it's where the actual tuning happens. You're looking for the calls where the agent technically succeeded but the caller sounded confused, because those don't show up in any success metric.
Widen the scope only when the escalation rate on that call type is boring. Then take the next call type and repeat.
We work this way with clients because it's the only version where problems surface while they're still small. It's also the reason our engagements are set up as ongoing tuning rather than a build and a handover. An agent that nobody reviews gets worse relative to the business around it, because the business keeps changing and the agent doesn't.
#The one line summary
An AI receptionist is a very good front door and a poor negotiator. Give it the volume, the repetition and the hours nobody wants to cover. Give a person the judgement, the money and the difficult feelings. Write the line between those two things down before you launch, not after a customer finds it.
If you want help drawing that line for your own call mix, tell us what your phones do on a normal Tuesday and we'll map which call types are worth handing over first.
Common questions
Still wondering
Will callers know they are talking to an AI receptionist?
Some will and some will not, and that matters less than how the call goes. What gives it away is usually latency rather than wording, because a pause of a second and a half reads as wrong in a way no phrasing can fix. We tell clients to disclose it plainly if asked, and never to have the agent claim to be a named human being.
What happens when the agent cannot handle a call?
It escalates according to rules you set. That normally means a warm transfer to whoever is on call, a callback booked into a real diary slot, or a message taken with the full transcript attached. The important part is that the boundary is defined before launch rather than discovered by a customer at the worst possible moment.
Can an AI receptionist book directly into our calendar?
Yes, and it should. An agent that takes details but cannot commit a slot has only moved the work rather than removed it. Booking means live access to real availability, your own rules about job length and travel time, and a written record that the caller receives too. Without those three the booking is a suggestion, not an appointment.
How long does it take to get a voice agent live?
Most deployments run three to six weeks, and the variable is almost never the agent itself. It is how many systems have to be connected and how quickly someone internally can answer questions about pricing, availability rules and edge cases. A single call type on an existing calendar can be live much faster than that.


