Why your AI receptionist gets names wrong on calls

The phone network throws away the part of your voice that tells an S from an F, and no model gets it back.

An abstract wave composition showing a voice signal narrowing as it passes through a telephone band

Say you run a three van boiler repair firm in Sheffield. Friday, 7pm. A call comes in, your AI receptionist takes it, books a Monday slot and writes the job to your calendar. Monday, the engineer drives to 15 Thorncliffe Road. The boiler is at 50 Thorncliffe Road.

Nobody hallucinated anything. The agent heard "fifty" and wrote "fifteen", which is the oldest mistake in telephony, on a line that has been discarding the evidence since 1988.

An AI receptionist gets names wrong because the phone network carries only 300 to 3400 Hz, and most of what separates one letter or syllable from another in English sits above that. No speech model recovers audio that was never transmitted. The fix is not a better model, it's designing the call so that nothing expensive gets acted on without a readback or a keypad press.

Why does my AI receptionist get names wrong?

Because the line deletes the frequencies that tell similar sounds apart, and a surname is the one thing on the call with no context to fall back on.

ITU-T Recommendation G.711, still in force and still underneath most call paths, passes audio between 300 Hz and 3400 Hz and samples it at 8 kHz. The energy that distinguishes an /s/ from an /f/ lives well above 3400 Hz. So does a good deal of what separates a /t/ from a /p/.

For ordinary words this barely matters. A language model fills the gap from context, because "I need a boiler ser..." is only ever going to end one way. A surname has nothing to fill from. Shaw, Shore and Shah are all equally plausible, so the model picks whichever was most common in its training data. That is how a Nguyen becomes a Wynn.

House numbers fail the same way. "Fifteen" and "fifty" differ by one unstressed syllable at the end of a word, sitting in the band the line squeezes hardest.

Word error rate hides what an AI receptionist gets wrong

Vendors quote WER (word error rate, the share of words a speech model gets wrong across a whole transcript). That number is genuinely good now. It is also the wrong number.

A booking rests on about five tokens: surname, house number, mobile, postcode, slot. Get 98 percent of a call right and lose the postcode, and the job is 100 percent failed. Averages do not help you here.

The measure that catches this is entity error rate, sometimes written NEER (named entity error rate, the share of names, numbers and addresses that come out wrong). AssemblyAI publishes both figures for its own real time model on its current benchmarks page: a 6.99 percent pooled word error rate on the open Pipecat benchmark of real agent conversations, against a 15.31 percent entity error rate on the same kind of material. Names are the worst category at 16.92 percent. Phone numbers are the best at 3.55 percent, because digits have a tight grammar and the model knows how many to expect.

Those are a vendor's own numbers, published to win a comparison. Even so, the entity figure is more than double the word figure, and surnames fail roughly one call in six. That gap is the distance between a demo and a Monday morning.

Does spelling the name out loud help an AI receptionist?

Usually less than you think, and sometimes it makes things worse.

Spelled letters are the hardest audio on a telephone line. English letter names cluster into what speech researchers call the E-set: B, C, D, E, G, P, T, V and Z. They share a vowel and differ only in a short burst at the front, which is exactly the part a 300 to 3400 Hz passband mangles. A caller spelling "Beattie" hands the agent B, T and E back to back. You've made them repeat themselves and handed the recognizer a harder problem than the one it just failed.

Two things do work.

The first is the ICAO spelling alphabet, the Alfa Bravo Charlie set that the International Civil Aviation Organization finalised in 1956 and the ITU later adopted. It exists because radio operators hit this bandwidth problem decades before anyone built a voice agent. "B as in Bravo" wraps a whole stressed syllable around a letter whose own distinguishing sound is a burst too brief and too high to survive the line. OpenAI's own realtime prompting guide tells builders to read characters back one at a time with separators and to add a phonetic disambiguator for letters. That is 1956 aviation practice, pasted into a system prompt.

The second is the keypad. DTMF (dual tone multi frequency, the tones your phone sends when you press a key) is not speech at all. It's two sine waves, decoded deterministically, and it survives a line that would destroy a spoken digit. Any agent taking a mobile number, a house number or a date of birth should offer the keypad up front, not as a consolation prize after it has misheard twice.

Which details should an AI receptionist confirm on every call?

The ones where a wrong value cancels the job: the mobile number, the address, the date and time.

A warped grid pulled out of shape by a single heavy point, showing how a few intake fields carry nearly all the risk on a call.
A warped grid pulled out of shape by a single heavy point, showing how a few intake fields carry nearly all the risk on a call.

Confirmation isn't free. Every readback is a sentence from the agent and one back from the caller, so an agent confirming nine fields is an agent people hang up on. Budget it. We sort intake fields into three groups before writing a single prompt:

  1. Confirm always. Mobile number, house number, postcode or building name, appointment date and time. Read digits back in pairs, and attach the day name to the date so "the fourteenth" becomes "Tuesday the fourteenth".
  2. Confirm by exception. Surname, company name, email address. Confirm when recognition confidence is low, or when the value matches nothing already in your CRM. A surname that's already on file against that mobile number doesn't need spelling out again.
  3. Never confirm. The free text description of the fault, how they found you, whether the dog is friendly. Getting "combi boiler" slightly wrong costs nothing, and a human reads it before anyone drives anywhere.

One rule matters more than it sounds: corrections have to propagate. When a caller fixes the fourth digit of their mobile, the agent repeats the whole corrected number, not the fix on its own. We've seen agents accept a correction, acknowledge it warmly, then write the original value anyway, because the tool call was built from the first turn rather than the last. Nothing errors. The record is just wrong, which is the same silent failure as an automation that quietly stops firing.

How do I know if my AI receptionist is getting names wrong?

Count it. The reason this runs for months is that nobody measures it.

Pull last month's calls that ended in a booking. For each one, line up three things:

  • What the caller actually said, from the recording.
  • What the agent read back to them.
  • What landed in your CRM.

Any disagreement between those three is a defect, even on calls where the job went fine. If you aren't certain the CRM end is wired up at all, settle that first: we've written about how to tell whether a CRM integration is real.

You want an error rate per field, not one overall score. One or two fields carry nearly all the damage, and they're reliably the ones nobody thought to confirm. Four wrong mobile numbers in fifty bookings is eight percent of your pipeline you physically cannot ring back, which makes every argument about how fast you respond to a lead academic.

Where an AI receptionist is the wrong tool

If your first call needs more than about six captured fields, don't buy one, and we'll tell you that before you sign anything.

Insurance intake, clinical triage, anything built on a twelve field form: confirm every field and the call runs long enough that people hang up, confirm nothing and the record is untrustworthy. No setting fixes that. We capture the three fields that let a human ring back, send an SMS with a form link part way through the call, and let the caller finish on a screen where they can see what they typed. Smaller job for us, better outcome for them.

The second limitation is ours, not physics. We have shipped agents with readbacks on the obvious numeric fields and none on the surname, because the surname tested fine in a quiet room, and gone back afterwards to add them. Nobody can make a model hear a letter the network never carried. Getting the field list right first time is on us, and we have not always managed it.

What changes for voice agent accuracy in the next two years

The pipe gets wider, slowly.

On 28 October 2025 the FCC, the US telecoms regulator, adopted a notice of proposed rulemaking, FCC 25-73. It proposes to drop incumbent carrier obligations to keep interconnecting over TDM (time division multiplexing, the circuit switched plumbing of the old phone network) by 31 December 2028, moving voice interconnection to IP. That rulemaking covers interconnection duties, not audio quality, and it's a proposal rather than a rule. The direction still matters. Fewer legacy hops in a call path means more calls that can stay wideband end to end, and a wideband call hands a speech model the detail above 3400 Hz that tells an S from an F.

There's something to do about it now and it costs nothing. Next time you review your phone provider, ask whether they pass G.722 or Opus end to end on inbound calls, and whether your voice agent's telephony layer accepts wideband audio or down samples everything to 8 kHz on the way in. Plenty of stacks do the second by default and never mention it. That's a configuration setting, not a rebuild.

Start by counting

Take fifty booked calls from last month. Check the mobile number and the address on each one against the transcript, and write down the error rate per field. An hour of that tells you which two fields need a readback, and a readback is a prompt change rather than a new system.

If you'd rather we counted with you, send us what you've got running and we'll go through a sample of your transcripts together. It's the first thing we do on any voice agent we take over.

Common questions

Still wondering

Will a newer AI model stop my receptionist mishearing surnames?

Only a little. A better recognizer helps at the margins, but it cannot recover frequencies the telephone network never carried in the first place. The reliable gains come from call design: reading values back, offering the keypad for anything numeric, and checking a captured surname against records you already hold. Treat model upgrades as a small bonus rather than the fix you are waiting for.

Should the agent ask callers to use the phonetic alphabet?

Not as the opening move, because it sounds officious on a first call. Use it as the second attempt. When the agent has already asked once and the value still looks uncertain, switching to the ICAO set and saying something like B as in Bravo gives the recognizer far more signal per letter. Keep it to the one field that actually failed.

Are mobile calls worse than landline calls for voice agent accuracy?

Often yes, and for two separate reasons. Mobile codecs compress harder than a fixed line, and callers on mobiles are usually somewhere noisy: a van, a street, a site. Both effects land hardest on exactly the short consonant bursts that distinguish letters and the ends of numbers. If most of your inbound calls arrive from mobiles, your confirmation rules need to be stricter.

What should the agent do when it is not confident about a value?

Say so and ask again, rather than guessing and moving on. A good agent reads back what it thinks it heard, invites a correction, then repeats the full corrected value so the caller can hear the final version. If two attempts fail on a field that matters, it should offer the keypad or take a callback number and pass the call to a person.