ProductKiosk AIWebsite AIIndustriesUse CasesPricingBlogSecurityPartnersContact Request a Demo
Technical

Measuring Voice AI Accuracy: Beyond "It Sounds Smart"

A practical framework for voice AI accuracy evaluation — how to score grounding, resolution rate, unmet queries, and escalation before you buy.

You measure voice AI accuracy the way you'd audit any system you trust with customers: test whether answers trace back to your sources, count how many questions actually get resolved, inspect what happens to the ones that don't, and check how cleanly it hands off to a human. "It sounds smart" is a demo reaction, not a metric — and it's exactly the impression a fluent model gives right before it invents your opening hours.

Why "it sounds smart" is the trap

Language models are optimised to sound plausible, which means fluency and correctness are separate axes. An answer can be confident, well-phrased, and completely wrong. On a website the user can click a link and check; in a lobby the spoken answer is the whole interaction, with no footnote to correct it. So voice AI accuracy evaluation can't rely on vibes from a scripted demo. It needs a repeatable method that scores the behaviours that matter and, crucially, uses your content and your questions — not the vendor's rehearsed prompts.

A note on scope: this is about evaluating accuracy before and during a buying decision, using a controlled test set. It's a different job from the live operational dashboards you'll watch once you're in production — those are covered separately in what to measure in voice AI analytics. Here we're pressure-testing the system's honesty before you commit.

Metric 1: Grounding rate

Grounding is the foundation. For a sample of questions with known answers in your knowledge base, what share of responses are actually supported by a real source passage — and can the system show you which one? This is where retrieval-augmented generation earns its keep: a grounded system retrieves your source text and answers from it, rather than from the model's general memory.

To score it, take fifty to a hundred questions whose correct answers live in your documents, ask them, and mark each response as: grounded and correct, grounded but incomplete, or ungrounded (plausible but not traceable to a source). The ungrounded bucket is the dangerous one. A mature platform lets you inspect the retrieved passage behind each answer, so grounding becomes something you can audit rather than take on faith.

Metric 2: Resolution rate

Resolution rate is the headline number: of everything people asked, how many got a useful answer without hitting a dead end or needing a human. It's the metric that separates a system that's merely switched on from one that's genuinely helping. Track it overall and by topic — a low or falling rate in one category points straight at a content gap.

When you build your evaluation, resist the urge to test only the easy questions. Mix in the awkward phrasing, the regional accents, the half-finished sentences, and the multi-part questions your visitors actually ask. Resolution rate on a clean, English, well-articulated test set will flatter every vendor equally; resolution rate on realistic input is what predicts how the system behaves in your building.

Metric 3: Unmet-query handling

No knowledge base is complete, so the most revealing test is what happens when the answer genuinely isn't there. Deliberately ask questions your content doesn't cover and watch the response. The right behaviour is a clear "I don't have that — let me get someone who can help," never a confident fabrication. A system that invents an answer for a question it can't ground has failed the single most important test, no matter how good its resolution rate looks elsewhere.

Score this explicitly: for a set of out-of-scope questions, what share are handled with an honest refusal versus a plausible guess? This is refusal behaviour, and it's the difference between a grounded assistant and a confident guesser in a nice voice. As a bonus, every unmet query the system logs is a free instruction for what to add to your knowledge base next.

Metric 4: Escalation quality

Accuracy includes knowing your own limits. When the system escalates — for anything sensitive, unusual, or beyond its knowledge — measure two things: does it escalate at the right moments (not too eagerly, not too late), and does it hand off without dropping the thread, so the person doesn't have to start over? A clean escalation that carries context is part of a correct answer, not an admission of failure. Test the boundary cases specifically: the emotional visitor, the security-adjacent request, the question that's 80% answerable and 20% needs a human.

The four numbers that matter — grounding, resolution, honest refusal, and clean escalation — all measure the same underlying quality: does the system answer from your reality, and admit when it can't?

How to run the evaluation

Turn the four metrics into a repeatable test rather than a one-off impression:

  1. Build a test set from real questions. Pull the actual top questions your front line hears, plus a batch of deliberately out-of-scope ones. Include your real languages, not just English.
  2. Load your own content. Insist on a pilot grounded in your documents. A demo on the vendor's sample data tells you nothing about your deployment.
  3. Score blind where you can. Have someone who didn't write the questions mark each response against the four metrics, so fluency doesn't bias the grade.
  4. Review the transcripts. The failures are more instructive than the successes — each one is either a content gap or a behaviour to fix.

What good looks like (and the red flags)

Strong systems score high on grounding and refusal even when resolution rate has room to grow — because resolution rate climbs naturally as you fill content gaps, whereas a system that fabricates can't be fixed by adding documents. Be wary of any platform that resists using your content, can't show the source behind an answer, has no honest-refusal behaviour, or only performs on scripted prompts. Those are the signs of a system optimised to impress evaluators rather than serve visitors. The same rigour applies whether you're grounding a web or kiosk deployment — the interface changes, but the accuracy questions don't.

Takeaway: Judge voice AI accuracy on four measurable behaviours — grounding, resolution rate, honest handling of unmet queries, and clean escalation — tested with your own content and your visitors' real questions. Fluency is easy to fake; groundedness is what you're actually buying.

See Kuyil for yourself

A live, 15-minute conversation with your future front desk — in any language.

Request a Demo
Keep reading

Related articles

Designing a Voice Persona: TTS Choices That Build Trust

Learn how to design an AI voice persona and pick a TTS voice that builds trust by matching voice, tone, and pacing to each deployment environment.

Read article

Data Residency and Sovereignty for Voice AI

AI data residency vs data sovereignty for voice AI: what each term means, in-region and on-prem options, and exactly what to ask for in regulated geographies.

Read article

Integrating Voice AI: Directories, CRM, and Host Notifications

How voice AI integrates with your stack — Slack and Teams notifications, CRM and ticketing, SSO directories, and industry systems via secure APIs.

Read article
FAQ

Frequently asked questions

Voice-first AI greets, listens and answers out loud, working on kiosks and in physical spaces as well as the web — reaching people a text chatbot cannot.
It uses retrieval-augmented generation (RAG): answers are grounded in your own documents, with citations, and it escalates to a human when unsure.
Kuyil supports 50+ languages, with automatic detection and mid-conversation switching.
On voice kiosks in lobbies and public spaces, and as a voice + text assistant on your website — all from one shared knowledge base.
Yes — tenant isolation, encryption, configurable retention and audit trails, with SOC 2 / ISO 27001 posture and HIPAA-ready options.
Under a second, so conversations feel natural rather than laggy.