AI Receptionist: What Actually Breaks When You Run One
Every AI receptionist demo is a happy path. We ran two weeks of calls through our own voice agent and logged the unhappy ones. Almost a fifth ended in under 20 seconds, and not one of them produced an error.

Every AI receptionist demo you will see is a happy path. Someone calls, speaks clearly, asks for an appointment, and the agent books it. Thirty seconds, no friction, cut to the pricing page.
We built one and then went looking for the unhappy paths. Over two weeks in late September and early October 2026 we ran a block of inbound and outbound calls through our own voice agent on a Twilio number. This article is what we found: how those calls actually ended, the five failures that never appear in a demo, and the questions worth asking a vendor before you sign.
One disclosure up front, because it changes how you should read the proportions below. These were our own calls and invited testers, placed to exercise the system rather than to sell anything. They tell you how the technology behaves. It does not tell you what your conversion rate will be, and anyone quoting you a conversion rate from their own test calls is doing something we are deliberately not doing here.
What an AI receptionist actually is
An AI receptionist is software that answers your phone, holds a spoken conversation with the caller, and takes an action in one of your business systems: booking an appointment, capturing a lead, routing the call, or answering a question from a knowledge base.
Mechanically it is four parts in a loop. Speech to text turns the caller's audio into words. A language model decides what to say and which tool to call. Text to speech turns the reply back into audio. A telephony layer, usually Twilio or a reseller of it, carries the call. The whole round trip has to complete in under roughly a second, or the caller starts talking over the agent.
The important thing about that stack is where the failures live. Almost none of them are in the language model, which is the part the demos are showing you.
How the calls actually ended
Where the call lengths landed
- Under 20 secondsno conversation took place19%
- 20 to 59 secondspartial qualification32%
- 60 to 119 secondsfull qualification28%
- 120 seconds or longerbooking conversations21%
View as a table
| Call length | Share |
|---|---|
| Under 20 secondsno conversation took place | 19% |
| 20 to 59 secondspartial qualification | 32% |
| 60 to 119 secondsfull qualification | 28% |
| 120 seconds or longerbooking conversations | 21% |
The bar that matters is the first one. Almost a fifth of connected calls ended in under 20 seconds, which is not long enough to have had a conversation. Those are the calls where something went wrong: the agent did not start speaking, the caller heard silence and hung up, or the audio never arrived in the first place.
Here is the uncomfortable part. Not one of those calls produced an error. Twilio recorded them as completed. The agent process stayed up. Every dashboard was green. If we had been measuring uptime, we would have reported 100%.
Five failures a demo will not show you
1. A misheard email address never throws an error
This is the one that costs real money, because it fails silently and downstream.
Speech to text does not return an error when it mishears. It returns a plausible word with high confidence. An address dictated over the phone comes back as a valid-looking string, the agent writes it to the CRM, the calendar invitation is sent, and nobody finds out until the meeting time passes with an empty room.
We saw the same tester's first name transcribed three different ways across our test bookings. A name is cosmetic. An email address is the delivery address for the only artefact the call produces.
We ended up refusing to book on an address the caller has not confirmed. The agent reads the address back one letter at a time, waits for a yes, and the booking tool returns an error to the model if it is called before that confirmation exists. Ask any vendor what their agent does with an email address it is not sure about. "It asks them to repeat it" is not the same as "it will not book without a confirmation."
2. The name on the booking is whatever the system last heard
Our agent matches an inbound caller to an existing CRM contact by phone number, which is correct: someone calling back should land on their own record rather than create a duplicate. It then writes the name they give on this call to that record.
Chain those two behaviors together and the last caller from a given number silently renames the contact. On a shared office line, a reception desk, or a phone two people use, the appointment created for one person can carry another person's name. We found this by looking at a run of bookings on a single test number and watching the name change under us several times.
It is a small bug with an embarrassing failure mode, and it exists in any system that keys on phone number and trusts spoken input. Ask whether the agent overwrites CRM fields or appends to them.
3. Silence has no error code
When a voice agent stops talking, nothing anywhere reports a problem. The websocket is open, the call is connected, the process is running. The caller is listening to nothing.
We only caught this by measuring the gap between the agent finishing a sentence and the next audio frame. The intuitive measurement, which is when the agent starts speaking, is useless: a long answer and a dead session look the same at the start. You have to measure from the stop event.
4. Answering machine detection can eat your opening line
Twilio's answering machine detection can run synchronously, in which case the call is held while it decides whether a human or a voicemail picked up. That hold happens during the first seconds of the call, which is exactly when your agent is delivering its opening line. The caller says hello, hears nothing, and hangs up.
Running the detection asynchronously fixes it. This is a configuration flag, not a product limitation, which is why it is worth asking about: a vendor who has not hit it has not run enough outbound calls.
5. The calling window encodes a timezone assumption
Outbound agents need a window of hours when it is acceptable to ring someone. Ours derives the caller's timezone from their phone number, falls back to a conservative window when it cannot, and refuses to dial outside it.
Then a UK enquiry arrived and the dialler would not ring it. The number mapping only covered North American area codes, so a +44 number fell through to a New York fallback and was held until 17:00 London time. The lead had submitted the form in the middle of their own working morning.
The guardrail was working exactly as designed. The design was wrong. If you buy an agent that calls people, ask what it does with a number from a country it does not recognize.
What to ask before you buy
| Ask this | Because | A weak answer sounds like |
|---|---|---|
| What happens to an email address the agent is unsure about? | Misheard addresses never error and break silently after the call | It asks the caller to repeat it |
| How do I find out the agent stopped speaking mid-call? | Silence produces no error anywhere in the stack | You can listen to the recordings |
| Does it overwrite CRM fields or append to them? | Phone-number matching plus spoken input renames contacts | It keeps the CRM in sync |
| What does it do with a phone number from a country it does not recognize? | Calling windows encode a default timezone | It uses the lead's timezone |
| Show me the distribution of call lengths, not the average | Averages hide the short calls, which are the failures | Our average handle time is X |
| What is the alert when the whole thing stops answering? | A process that is gone cannot report that it is gone | We monitor it 24/7 |
That last row is worth dwelling on. Every check that runs on the same machine as the agent dies with the machine, and the silence looks exactly like health. The only honest answer involves something watching from outside, on separate infrastructure. We run ours from a scheduled job on a different provider entirely, for that reason.
Build or buy
Buy if you want a receptionist that answers, qualifies and books into a common calendar, and your business logic is ordinary. The products in this category do that well now, and you will not beat them on price or on time to launch.
Build if the agent needs to make decisions against your own data, write to systems with awkward APIs, or follow rules you cannot express in a vendor's form builder. Also build if the call is the product rather than a front door to it, because at that point the behavior of the agent is your competitive surface and you should not be renting it.
The hidden cost in both directions is the same, and it is not the model tokens. It is the instrumentation: knowing a call went silent, knowing a booking carried a bad address, knowing the thing stopped answering at 2am. Budget for that work whichever way you go, because it is what separates an agent you can leave running from a demo you have to supervise.
The honest summary
The technology works. Roughly half our connected calls ran past a minute and a fifth ran past two, which is a real conversation by any standard, and the booking flow does write a correctly formed appointment into a CRM without a human touching it.
What is not solved is that this category fails quietly. Every failure above produced a call that Twilio logged as completed, a process that stayed up, and a dashboard that stayed green. If you deploy one of these, the work that actually determines whether it is an asset or a liability is the measurement you put around it.
Key takeaways
- Ask a vendor for the distribution of call lengths, not the average. Averages hide the short calls, and the short calls are the failures.
- Ask what the agent does with an email address it is unsure about. Asking the caller to repeat it is not the same as refusing to book without a confirmation.
- Ask whether the agent overwrites CRM fields or appends to them, because phone-number matching plus spoken input silently renames contacts.
- Ask what happens to a phone number from a country the system does not recognize, since calling windows encode a default timezone.
- Ask where the monitoring runs. A process that has died cannot report that it has died, so the check has to live on separate infrastructure.
- Buy if your business logic is ordinary. Build if the agent needs your own data or the call is the product rather than a front door to it.
Questions people ask about this
Sourced from client conversations, Search Console, and AI-search citation monitoring.
- Software that answers your phone, holds a spoken conversation with the caller, and takes an action in one of your business systems: booking an appointment, capturing a lead, routing the call, or answering from a knowledge base. Mechanically it is speech to text, a language model that decides what to say and which tool to call, text to speech, and a telephony layer, with the whole round trip completing in roughly a second.
- Less than the marketing suggests. The language model is rarely where deployments fail. The failures sit in the telephony configuration, the speech recognition confidence, the CRM write path and the monitoring, which is to say in ordinary software engineering rather than in the model.
- A call that connects and produces no conversation. In our own testing, almost a fifth of connected calls ended in under 20 seconds. The telephony provider recorded all of them as completed, the agent process stayed up, and no dashboard showed a problem. Uptime monitoring would have reported 100%.
- Yes, and it is the failure that costs the most because it surfaces latest. Speech to text does not raise an error when it mishears; it returns a plausible string with high confidence. The agent writes it to the CRM, the invitation goes to a nonexistent address, and nobody finds out until the meeting time passes. The defense is a spelled-back confirmation that the booking tool refuses to proceed without.
- Buy if you want answering, qualification and booking into a common calendar and your business logic is ordinary, because you will not beat the products on price or time to launch. Build if the agent must decide against your own data, write to systems with awkward APIs, or if the call is the product rather than a front door to it.
- The per-minute cost of speech, model and telephony is the small part and it is published by every vendor. The cost people underestimate is instrumentation: knowing a call went silent, knowing a booking carried a bad address, knowing the system stopped answering overnight. Budget for that work whether you build or buy, because it is what makes the agent safe to leave running.
- It can, if the calling window derives a timezone from the phone number and the number is from a country the mapping does not cover. Ours held a UK enquiry until the London evening because a plus 44 number fell through to a North American fallback. The guardrail worked as designed and the design was wrong, which is the normal shape of this bug.
- Only if something outside it is checking. Every monitor that runs on the same machine as the agent stops when the machine does, and total silence is indistinguishable from a quiet day. Run the check from separate infrastructure, and ask to see it go red on purpose before you trust it.
Keep reading
Related articles.
More on the same thread, picked by tag and category, not chronology.
AI Agents for Sales: Lead Qualification to Close
AI agents automate the sales cycle from lead qualification through close. Learn how sales teams use AI agents to 3x pipeline velocity in 2026.
The AI Agent Tech Stack: What You Need to Build Production Agents
The complete AI agent tech stack for production deployments. LLMs, frameworks, memory, tools, observability, and guardrails, everything you need in 2026.
AI Agents vs Chatbots vs RPA: Understanding the Differences
AI agents, chatbots, and RPA serve different automation needs. Understand the key differences in capabilities, costs, and use cases to choose the right technology.
Ready when you are
Want Cubitrek to run AI Agents for you?
We install ai agents programs for growing companies across the US and Europe. Book a call and we'll come back with a one-page plan in 72 hours.

