Fail-safe tools for a voice AI receptionist
By Matthew Gray on Oct 6, 2026
A voice agent’s tools fail in ways a chat agent’s don’t: the caller hangs up mid-sentence, a setting is missing, or the phone network refuses a transfer while someone is listening. The lesson from building our own AI receptionist is that every tool needs a safe fallback that still captures the caller, and that telephony features need proof on live calls, not just passing unit tests.
The receptionist answers Hat Boy Software’s main line. It answers basic questions about us, records new callers as leads in our CRM, opens support cases for existing clients, and checks office hours and holidays. It runs on Azure Communication Services for the phone number and call control, Azure VoiceLive with a realtime voice model for speech, and Azure Container Apps for the backend. We started from Microsoft’s open-source, MIT-licensed voice agent accelerator rather than writing the audio pipeline ourselves. It is still evolving, and this post includes what isn’t finished.
Every tool must degrade safely
On a phone call there is no retry button; a failed tool sounds like a pause, or nothing. So each tool follows one rule: whatever goes wrong, the caller’s request still reaches a person.
| Failure | What happens |
|---|---|
| Caller hangs up mid-intake | On shutdown, the call handler posts a lead with the caller ID and the call transcript |
| Transfer requested, no number configured | The tool returns “not configured” and the agent takes an urgent message |
| Transfer fails or the number is bad | Same: the agent apologises and takes a message, without retrying |
| No support-case token configured | The request is recorded as a lead, tagged with its urgency |
| Client has no support agreement | A lead is recorded, not a case, with the scope of what they asked for |
The hang-up fallback took the most care. If the caller spoke but the agent never submitted a lead, the handler posts one anyway, named “Unknown caller”, with the transcript in the notes. The placeholder name exists because our CRM rejects leads with no name, email or company. Calls with no caller speech, or no caller ID, are skipped. The transcript is trimmed to stay under the CRM’s 4,000-character notes limit.
Tool results also tell the model what to do next, in plain words. A failed transfer doesn’t return a bare success: false. It returns that, plus an instruction to take a message and say the team will be alerted. The prompt then says not to try again. Left with a raw error, a model may retry or invent a reason.
Recognising callers without widening access
Existing clients shouldn’t have to spell their email address to a robot. When the caller first speaks, a lookup tool matches the caller ID against client phone numbers in the CRM. The match is cached for the rest of the call. A recognised client is greeted by name and company, and the request goes to the right place for their support tier.
- Compare the last 10 digits. Phone numbers arrive with or without a country code, with brackets, dashes or spaces. We strip everything but digits and compare the last ten, so the same number in two formats matches. Anything shorter than ten digits doesn’t match at all.
- Use a separate, read-scoped key. Creating a lead uses a public web-form token, the same kind our website’s contact form embeds, which can only create leads. Looking up clients needs read access to client records, so it uses a different key. That key is read-scoped and kept as a secret in Key Vault. If it is missing, nobody is recognised and every caller is treated as new, which is the safe direction to fail.
Normalise messy configuration
Support tiers come from a custom field on each client record, and settings are typed by people. “Tier 1”, “tier-1”, “TIER1” and “1” should all mean the same thing. Rather than make staff type a value exactly right, the code normalises every tier name before it is used:
def tier_key(tier):
key = "".join(ch for ch in str(tier or "").lower() if ch.isalnum())
return f"tier{key}" if key.isdigit() else key
A tier that still isn’t recognised falls back to the default, lowest tier and logs a warning, rather than failing the case. Staff can re-tier it later. Malformed JSON in a settings map is logged and ignored, not fatal at startup.
Tune the conversation for how people actually talk
- Silence threshold. We raised the end-of-turn silence threshold to 800 ms. Callers pause while spelling an email address, and the agent shouldn’t take that pause as the end of their turn.
- Read-back. The agent spells email addresses back letter by letter and corrects them until the caller confirms. It never guesses a spelling.
- Guardrails. It never quotes prices, rates or estimates. If asked whether it is a person, it says plainly that it is an AI assistant that takes messages for the team. It never asks for payment details or passwords, and the greeting says the call may be recorded and transcribed.
Right-size the accelerator
The accelerator is built for enterprise contact-centre load. A small business main line is not that, so we cut it down:
- Replica floor of one. The backend runs one to two replicas at 1 vCPU and 2 GiB, against an upstream default minimum of five replicas at 2 vCPU. We kept the floor at one, not zero, so calls are answered at once instead of waiting on a cold start while the phone rings.
- Remove what the receptionist doesn’t use. The document database, a demo service that depended on it, the demo models and the email service are all switched off. The backend starts without the database; only demo features and some analytics stop working.
- Check availability region by region. Not every SKU and model is offered in every region. In our deployment the realtime voice resources and the cache sit in different regions. Price the cache, now our largest fixed cost, for the region you choose before going live.
Verify what unit tests can’t prove
Live transfer is the clearest example. The transfer tool calls the telephony service’s transfer API, and it has unit tests, but those tests stub that call. They prove our routing and fallback logic. They can’t prove that a real inbound call reaches the destination, that audio flows both ways, or that the receptionist stops talking when the call bridges.
So transfer shipped switched off. With no destination number set, the agent takes a message, exactly as before. Before we relied on it, our runbook required:
- two consecutive clean live transfers on a deployed environment, with the expected caller ID at the destination and no errors in the logs
- both fallback paths checked on a live call: no number configured, and an unreachable number
- a caller hanging up mid-transfer, with the end-of-call fallback still firing
We have now run that verification on live calls. Transfer stays off by default and is turned on per deployment by setting a destination number. Rollback is the same configuration change in reverse: unset the number.
Known gaps
- No warm transfer. Transfers are cold. The tool captures a summary for the person receiving the call, but nothing speaks it to them yet.
- It can’t hang up. The caller ends the call. The accelerator doesn’t act on an end-call result yet.
- Floating holidays are manual. Fixed-date holidays are built in. Holidays that move each year have to be added to the settings by hand.
How we can help
If you are considering an AI receptionist, the hard parts are the ones above: what happens when a tool fails, and proving telephony behaviour on real calls. We run ours on our own line first. You can read how we offer it on our AI receptionist page, or call our main line and talk to it.