We run our voice agent on Retell, on a number that stays ours. The criteria that decided it, and how to test a voice platform yourself.
We run our voice agent on Retell, on a number that stays ours. The criteria that decided it, and how to test a voice platform yourself.
We run our voice agent on Retell, over a carrier layer that abstracts Twilio and Telnyx so the phone number stays ours rather than the vendor's. This post is about how that choice got made, what we would test if we were choosing again tomorrow, and which numbers we are not willing to give you.
Start with the correction, because it is why this reads differently to the version you may have seen. An earlier one published a six week benchmark: error rates to one decimal place across four platforms, latency medians and tails, a fifty call accent panel, twenty blind listeners, five hundred webhook tests, and a fleet of live clients standing behind it. That benchmark did not happen. No panel, no listener group, no client fleet. We have rewritten the post rather than quietly pulling it, because the question underneath it is the right one. What follows contains no invented numbers.
Picture the call we build against. A bloke rings from a farm outside Mildura at 6:47 on a Tuesday morning. Broad accent, reception coming and going, a diesel pump running behind him. He needs an emergency plumber and he has already tried three numbers. He is not a customer and he is not a case study. He is an imaginary worst call, useful precisely because he is hard. Cope with him and you will cope with a specialist clinic in South Yarra at lunchtime. The failure you are designing away from is easy to picture too: the agent says "I didn't catch that" six times, then hangs up on him.
Why does voice AI stack choice matter more than most people think?
Most businesses treat voice AI like choosing another SaaS tool. Tick the boxes, pick the cheapest, move on. The stack underneath decides whether the agent answering your phone sounds like a person who wants to help or a GPS unit that got lost in 2014.
Three things go wrong when you pick the wrong one. Latency creeps up, and once the silence gets long enough the caller and the agent start talking over each other. Mishearing fails quietly: the agent believes it heard something sensible, carries on, and books the wrong job on the wrong day. And the write drops. The agent says "I've booked that in" and your CRM never hears about it. You find out three days later when they ring back, annoyed.
Pitcher Partners surveyed 316 businesses and reported in August 2026 that 90 per cent of finance and property firms say clients now demand faster responses. That is one sector, and it is the sector we know best, but the direction of travel is not in dispute.
Most AI voice tools are rubbish for Australian accents straight out of the box. The platform you pick is what decides whether you can do anything about that or simply live with it.
What criteria actually matter when comparing Bland AI, Vapi, Retell and Air AI?
We ended up on Retell, run through our own carrier layer rather than a bundled number, for three reasons. It lets us configure recognition and vocabulary hints for a given business rather than living with a single locked model, which matters most for suburb names and trade terms. It ships the handful of tool calls the agent actually needs, checking availability, booking an appointment, requesting a callback, and ending the call cleanly, without us having to build that infrastructure ourselves. And because the carrier layer sits between the platform and the number, that number stays ours whether a call routes over Twilio or Telnyx, so outgrowing a platform later does not mean losing the number a business already promotes.
Seven, in the order we would weight them. Buy a month of credit on the shortlist, point each one at the kind of call you actually get, and score them yourself.
- Latency. Time from the caller finishing a sentence to the first audible word back. Measure it off your own recordings with a stopwatch, not off a landing page.
- Australian accent recognition. Ring it yourself, from where your callers ring from, with the reception they have.
- Escalation reliability. When the agent cannot help, does the handoff land? A callback request that never reaches a human is worse than no agent at all.
- Write delivery. Trigger a CRM write on every test call and reconcile afterwards. Count the ones that needed a human to notice.
- Voice naturalness. Play a recording to someone who does not work for you and watch their face in the first ten seconds.
- Interruption handling. Talk over it. A caller with a toddler in the background will.
- Cost per minute at your volume. Real pricing under your call pattern, not the sticker.
We gave no weight to "emotional intelligence" or "enterprise grade". The only question is whether the phone works when somebody ready to spend money rings at 7am on a Monday.
How should you test Australian regional accents before you commit?
Here is where the old version handed you a table of error rates by vendor. We are not republishing it, and be sceptical of anyone who hands you one. A recognition score means nothing without the audio it came from, the sample size, and who scored it.
What we can give you is the method, which takes an afternoon. A worked illustration, not a result we are claiming:
Pick five voices from the places your callers actually live. Somewhere in regional Queensland, somewhere in Western Australia, one broad, one fast, one softly spoken. Four calls each is twenty calls, and across them the agent will take perhaps eighty conversational turns. Count the turns where it had to ask for a repeat. Three out of eighty is a product you can put on your main line. Twelve out of eighty will cost you callers. You do not need a decimal place to tell those apart.
Platform choice matters here for control rather than accuracy. The recognition model belongs to a vendor with a research budget nobody our size has, and we did not train it. What a good platform lets you do is feed that model the words your business uses: suburb names, trade vocabulary, the difference between a top-up and a refinance, what "bulk-billed" means. Lock that away and you are stuck with whatever shipped.
What latency thresholds separate a natural conversation from an awkward one?
These are the bands we build to. They are our own design targets rather than published research, and we are labelling them that way so you can argue with them.
- Under 600ms. Feels like talking to a person. Most callers do not clock it as software in the first half minute.
- 600 to 900ms. Noticeable, tolerable. Like a receptionist who is half awake.
- 900 to 1,500ms. Awkward. Callers start repeating themselves into the gap.
- Over 1,500ms. Broken. Callers assume the line dropped, then hang up or talk straight over the top.
The number people quote is the median, and the median is the least interesting one. Suppose a platform sits at a comfortable 700ms most of the time but one call in twenty stretches to 1.1 seconds. That is five calls in every hundred that feel wrong, and you do not get to choose which five. Ask for the tail, not the average. If nobody will give you the tail, measure it off your own recordings, because the tail is where your reputation gets spent.
What we run, and what we watch
The honest version of a proof section. theautomate is a Melbourne automation shop, and what we have built and operate is our own voice and intake platform. That is the only thing we can point at. No logos, no call counts, no answer rate to quote. Running your own software in production is a harder test than a case study anyway, because there is nobody to blame and no account manager to ring.
What we watch every week, in this order:
- Did the call get answered. Including the ones outside office hours, which are most of the interesting ones.
- Did the record land. The call happened and the CRM knows, with no human spotting the gap.
- Did the follow up fire. The message goes out without anyone remembering to send it.
- Did the booking land where it should. Right service, right slot, right day.
We will publish numbers when we have numbers worth publishing, and they will be ours. Until then the honest claim is about what the plumbing does, not what it achieved for somebody else. The same argument runs through what we have written on missed calls in Australian small business and the 2026 shift in the Australian AI receptionist market: the hard part was never the voice.
Honest limitations
Where the stack still struggles. Better you read it here than find it on a live call.
Very noisy environments. Cafes at peak, construction sites, a motorbike idling, strong wind. Noise rejection is good and not perfect, so we wire in a "couldn't hear you clearly, is there a better number to try?" fallback rather than pretending otherwise.
Rapid code switching. Callers moving between English and Mandarin or Vietnamese mid sentence throw the transcription off. Doing that properly needs a multilingual recogniser and a routing layer we have not shipped.
Legacy PBX. An on premise system with no SIP trunk needs hardware on site, so it is the first thing worth checking.
The first days of a new deployment. Edge cases in a booking flow only show up in real traffic, so go live is where the tuning starts, not where it stops.
Ours is a catalogue voice, not one we recorded. The picker inside the product says so in plain words, previews included, so you hear the provider's own sample before you commit to one. Several Australian voices sit in that picker. What we will not do is promise that any catalogue voice passes as a lifelong local: judge it by ringing the number yourself and listening, the same way you would judge anyone else's.
If a vendor tells you none of these exist on their platform, they are selling you the demo rather than the phone line.
FAQ
Can you share the raw benchmark data from your stack comparison? No, and the honest reason is that it does not exist. The spreadsheet the earlier version offered you was never real. What is above is the whole of what we have: the criteria, the method, and the thresholds we build to. If we ever run a comparison worth defending in public, the audio and the scoring get published with it.
Why not build your own voice stack from scratch? Because streaming speech that sounds natural is a full time engineering team for a year, and the layer that decides whether a call goes well is elsewhere: the script, the CRM wiring, the business rules, what happens when the agent gets stuck. The platform handles the hard telephony. We handle the part where the agent knows what a Tuesday afternoon appointment means for a one van plumbing business.
How often do you re-evaluate the stack as the market evolves? When something moves that touches one of the seven criteria. A platform shipping configurable recognition, a carrier changing its terms, a new entrant with a genuinely different latency profile. We are not married to a vendor, and the carrier layer exists partly so switching stays an engineering decision. The number stays ours either way.
What is the biggest technical risk in running AI voice agents at scale? Silent failures. Not outages, which show up on a dashboard within a minute. The quiet ones are where the agent believes it booked the appointment, the caller believes it too, and the CRM never got told. The defence is boring and it is the whole job: retry the write, reconcile the call log against what the CRM holds, then look at the mismatches by hand.
Does the platform choice affect the quality of Australian accent recognition? Yes, though not the way it is usually sold. The recognition model belongs to the vendor and you are not going to improve it. What changes between platforms is how much you can configure around it: the language setting, vocabulary hints for the words your callers say, how long the agent waits before deciding it missed something, and where it goes when it is stuck. With none of those on offer, broad Australian callers will usually leave you short.
Book 30 minutes with me
Book 30 minutes with me. I will tell you honestly whether this makes sense for your business, including when it does not. theautomate.io
Frequently Asked Questions
Can you share the raw benchmark data from your stack comparison?
Why not build your own voice stack from scratch?
How often do you re-evaluate the stack as the market evolves?
What is the biggest technical risk in running AI voice agents at scale?
Does the platform choice affect the quality of Australian accent recognition?
Written by Syed Bilgrami
Runs TheAutomate, a Melbourne automation agency. He scopes the work, writes it, and picks up when it breaks.
What is the work that repeats in your business?
Book a 30 minute discovery call with Syed. He scopes the work himself, so you are talking to the person who would build it.
Book a Discovery Call

