Designing a Voice Persona: TTS Choices That Build Trust
Learn how to design an AI voice persona and pick a TTS voice that builds trust by matching voice, tone, and pacing to each deployment environment.
Read articleA practical framework for voice AI accuracy evaluation — how to score grounding, resolution rate, unmet queries, and escalation before you buy.
You measure voice AI accuracy the way you'd audit any system you trust with customers: test whether answers trace back to your sources, count how many questions actually get resolved, inspect what happens to the ones that don't, and check how cleanly it hands off to a human. "It sounds smart" is a demo reaction, not a metric — and it's exactly the impression a fluent model gives right before it invents your opening hours.
Language models are optimised to sound plausible, which means fluency and correctness are separate axes. An answer can be confident, well-phrased, and completely wrong. On a website the user can click a link and check; in a lobby the spoken answer is the whole interaction, with no footnote to correct it. So voice AI accuracy evaluation can't rely on vibes from a scripted demo. It needs a repeatable method that scores the behaviours that matter and, crucially, uses your content and your questions — not the vendor's rehearsed prompts.
A note on scope: this is about evaluating accuracy before and during a buying decision, using a controlled test set. It's a different job from the live operational dashboards you'll watch once you're in production — those are covered separately in what to measure in voice AI analytics. Here we're pressure-testing the system's honesty before you commit.
Grounding is the foundation. For a sample of questions with known answers in your knowledge base, what share of responses are actually supported by a real source passage — and can the system show you which one? This is where retrieval-augmented generation earns its keep: a grounded system retrieves your source text and answers from it, rather than from the model's general memory.
To score it, take fifty to a hundred questions whose correct answers live in your documents, ask them, and mark each response as: grounded and correct, grounded but incomplete, or ungrounded (plausible but not traceable to a source). The ungrounded bucket is the dangerous one. A mature platform lets you inspect the retrieved passage behind each answer, so grounding becomes something you can audit rather than take on faith.
Resolution rate is the headline number: of everything people asked, how many got a useful answer without hitting a dead end or needing a human. It's the metric that separates a system that's merely switched on from one that's genuinely helping. Track it overall and by topic — a low or falling rate in one category points straight at a content gap.
When you build your evaluation, resist the urge to test only the easy questions. Mix in the awkward phrasing, the regional accents, the half-finished sentences, and the multi-part questions your visitors actually ask. Resolution rate on a clean, English, well-articulated test set will flatter every vendor equally; resolution rate on realistic input is what predicts how the system behaves in your building.
No knowledge base is complete, so the most revealing test is what happens when the answer genuinely isn't there. Deliberately ask questions your content doesn't cover and watch the response. The right behaviour is a clear "I don't have that — let me get someone who can help," never a confident fabrication. A system that invents an answer for a question it can't ground has failed the single most important test, no matter how good its resolution rate looks elsewhere.
Score this explicitly: for a set of out-of-scope questions, what share are handled with an honest refusal versus a plausible guess? This is refusal behaviour, and it's the difference between a grounded assistant and a confident guesser in a nice voice. As a bonus, every unmet query the system logs is a free instruction for what to add to your knowledge base next.
Accuracy includes knowing your own limits. When the system escalates — for anything sensitive, unusual, or beyond its knowledge — measure two things: does it escalate at the right moments (not too eagerly, not too late), and does it hand off without dropping the thread, so the person doesn't have to start over? A clean escalation that carries context is part of a correct answer, not an admission of failure. Test the boundary cases specifically: the emotional visitor, the security-adjacent request, the question that's 80% answerable and 20% needs a human.
The four numbers that matter — grounding, resolution, honest refusal, and clean escalation — all measure the same underlying quality: does the system answer from your reality, and admit when it can't?
Turn the four metrics into a repeatable test rather than a one-off impression:
Strong systems score high on grounding and refusal even when resolution rate has room to grow — because resolution rate climbs naturally as you fill content gaps, whereas a system that fabricates can't be fixed by adding documents. Be wary of any platform that resists using your content, can't show the source behind an answer, has no honest-refusal behaviour, or only performs on scripted prompts. Those are the signs of a system optimised to impress evaluators rather than serve visitors. The same rigour applies whether you're grounding a web or kiosk deployment — the interface changes, but the accuracy questions don't.
A live, 15-minute conversation with your future front desk — in any language.
Request a DemoLearn how to design an AI voice persona and pick a TTS voice that builds trust by matching voice, tone, and pacing to each deployment environment.
Read articleAI data residency vs data sovereignty for voice AI: what each term means, in-region and on-prem options, and exactly what to ask for in regulated geographies.
Read articleHow voice AI integrates with your stack — Slack and Teams notifications, CRM and ticketing, SSO directories, and industry systems via secure APIs.
Read article