//

11 min read

//

How to Evaluate Healthcare Voice AI Agents

A practical framework to evaluate healthcare voice AI agents: realistic test scenarios, three levels of metrics, patient safety guardrails, and post-launch tracking.

Key takeaways

Test with the calls patients make. Simulate accents, medical terms, age, education level, noisy lines, and patients who ask several questions in one call. Clean demo calls hide the failures that show up in the first week of live traffic.

Measure at three levels. Track task completion first, then task-specific rubrics like shared decision-making, then component accuracy such as word error rate and correct tool use.

Build safety into the workflow. A system prompt alone leaves the riskiest steps exposed. Guard tool calls, data access, and web search against PHI leakage, tool misuse, and medical advice.

Keep evaluating after launch. Track why patients ask for a human, sentiment, duplicate requests, and repeat callers on every call.

Pilot one workflow first. Start with a high-volume, well-defined task like prescription refills, and define success metrics before the first call.

In a pilot at a primary care provider, Level AI's AI Virtual Agent handled 45% of prescription refill calls without agent intervention. Refills were the highest-volume call type in that contact center, and the old IVR had sent every one of them to a clinical agent.

That 45% came out of evaluation work done before and after go-live. The team tested how the agent verified patients, watched where calls broke down, and changed the workflow each time the data pointed to a problem. The full story of how a healthcare provider automated 45% of prescription refill calls walks through each change.

Building a voice agent now takes days. Proving that agent is safe and effective for patients takes far longer, because a single healthcare call carries medication names, protected health information (PHI), and symptoms that belong with a nurse.

Evaluation breaks down in three places. Test scenarios rarely match how patients talk. Standard metrics score the agent's wording and ignore whether the patient's request got done. Safety checks live in a single prompt, far from the tool calls where data leaks. This guide covers each of the three, then what to track after launch, how to judge EHR integration, and what each leader on the buying team should check.

What Do Healthcare Teams Ask Before Adopting a Voice AI Agent?

Healthcare buyers bring the same six concerns to nearly every evaluation. Each one maps to a test the agent has to pass before it takes a live call.

1. Several questions stacked into one call

Patients rarely call about one thing. A single call might include refills for two medications, a question about a lab result, and a request to move next week's appointment. The agent has to track each request, complete the ones in scope, and route the rest without dropping any of them.

2. Urgent phrases that need escalation

When a patient says "I'm having a reaction" or "my chest feels tight," the call is no longer a refill request. The agent must catch the phrase, stop the current task, and transfer the patient to clinical staff with the call context attached. Buyers want proof that this escalation fires every time, including when the phrase shows up mid-sentence or halfway through another task.

3. Duplicate encounters and requests

A second telephone encounter for a refill already in progress adds work for clinical staff and delays the patient's medication. In the primary care pilot, duplicate documentation appeared when Salesforce did not surface existing telephone encounters. Any evaluation needs test calls from patients who already have an open request.

4. Background noise

Patients call from cars, pharmacy lines, and kitchens with the TV on. Noise raises transcription errors, and one misheard medication name sends the whole call down the wrong path. Buyers ask whether the agent recognizes that noise is the problem and tells the caller, so the patient can move somewhere quieter before the agent asks the same question a third time.

5. Speech-to-speech vs. speech-to-text-to-speech

A speech-to-text-to-speech pipeline transcribes the caller, generates a reply from the text, and converts that reply back into audio. Each step adds delay, and patients notice a pause longer than a second or two. Speech-to-speech models work on the audio directly and cut that delay. The tradeoff is control: the text pipeline produces a transcript at every step, which gives the team a place to run safety checks, log decisions, and audit the call later.

Ask vendors for latency measured on live calls, and ask where their guardrails run. Level AI's parallel multi-modal architecture for virtual agents runs safety checks, intent detection, and record retrieval side by side in about 300 milliseconds, then gives the reply its own 600 to 800 millisecond budget. The transcript layer stays in place, and the steps no longer wait on each other.

6. Whether patients will accept talking to an AI

Pilot data answers this better than any survey. Buyers want to see how many patients stay on the line, how many ask for a human in the first 10 seconds, and how sentiment moves from the start of the call to the end.

How Do You Build Realistic Test Scenarios When Healthcare Data Is Scarce?

Voice agents are judged on scenarios. A scenario is a full call: a patient with a medical history, a reason for calling, a way of speaking, and a set of conditions on the line. One correct answer to one question says little about how the agent handles the ten minutes of conversation around it.

Public healthcare datasets cover part of this. Synthea generates synthetic patient records. MIMIC-IV holds deidentified records from intensive care patients. MedAlign pairs clinician-written instructions with EHR data. None of them captures how your patient population talks on the phone, which medications they take, or how their calls to your contact center unfold from greeting to hang-up.

1. Variables to simulate

Build each scenario from a combination of these variables, and report results for each group separately so one group's errors do not disappear inside an average.

  • Accent and dialect: regional accents and non-native English speakers.

  • Medical terminology: brand and generic drug names, dosages, and common mispronunciations.

  • Education level: patients who describe symptoms in plain words and patients who use clinical terms.

  • Age: older callers who pause longer, speak more slowly, or ask for repetition.

  • Line quality: speakerphones, cell calls with dropouts, and background noise at several levels.

  • Call structure: several requests in one call, topic changes, and interruptions.

2. Grounded simulations

Grounded simulations combine three sources. Recorded patient calls and transcripts show how patients phrase requests and where calls go wrong. Interviews with front-desk and clinical staff surface the edge cases they handle every week. A language model then generates variations on those patterns, so one recorded call turns into 50 test calls across accents, noise levels, and request combinations.

Cover the normal path first. Add edge cases next. Finish with adversarial calls, where a caller tries to get information or actions they are not entitled to.

3. An edge case from the refill pilot

In the primary care pilot, a patient asked for a refill of "lipid drill." The patient meant lisinopril. An exact-match lookup against the medication list fails on that input, so the team evaluated fuzzy matching, which compares a misheard term against the patient's active medications and returns the closest match for the patient to confirm.

A scenario library needs dozens of cases like this one. The fastest source is the first few weeks of live calls, fed back into the test set before the next release.

Which Metrics Tell You If a Voice Agent Is Working?

BLEU, ROUGE, and BERTScore came out of language research, and they miss what matters on a patient call. BLEU and ROUGE count how many words the agent's reply shares with a reference answer. BERTScore measures how close the meaning is. A reply that confirms the wrong pharmacy scores well on all three when the wording is close, and none of them checks whether the refill was submitted.

Custom rubrics for every scenario fail in the opposite direction. A team with 400 scenarios and a unique rubric for each one spends its time maintaining rubrics, and scores from different rubrics cannot be compared release to release.

A three-level structure keeps scores comparable and points straight to the cause of a failure.

1. Task level: completion and efficiency

Task completion measures whether the patient's request ended in the correct outcome. That means the refill went to the right pharmacy, the appointment landed with the right provider, or the call reached the right queue. Efficiency measures the turns and seconds it took to get there. A refill that takes 12 turns still completes, and the patient on the other end notices every one of them.

2. Rubric level: task-specific checks

Rubric checks apply to a type of task and get reused across every scenario of that type. A shared decision-making rubric for scheduling checks that the agent offered the available options, confirmed the patient's preference, and repeated the choice back before booking. A refill rubric checks that the agent confirmed the medication, dose, and pharmacy, and flagged patients who will run out before their next appointment. Ten to fifteen rubrics cover most of the workflows in a primary care contact center.

3. Component level: the parts that feed the task

  • Word error rate (WER): the share of spoken words the transcription got wrong, counted as substitutions, deletions, and insertions divided by the total words spoken. Track it by accent, age group, and noise level, and track medication names as their own category. Accurate automatic speech recognition sets the ceiling for every metric above it.

  • Retrieval precision and recall: precision is the share of retrieved records that were relevant. Recall is the share of relevant records the agent retrieved. On a refill call, a recall failure means the agent missed one of the patient's active medications.

  • Tool-call accuracy: the right tool, with the right inputs, in the right order. Submitting a refill before identity verification counts as a tool-call failure even when the refill itself is correct.

When task completion drops, the component metrics show which step caused it. Scoring every scenario against these shared metrics on each release catches a regression before patients hit it.

Map buyer concerns to metric levels

Buyer concern

Metric level

What to measure

Several questions in one call

Task

Completion rate for each request inside multi-part calls

Urgent phrases

Rubric

Escalation triggered, time to transfer, context passed to staff

Duplicate requests

Component

Tool-call accuracy on existing-record lookups before any new record is created

Background noise

Component

WER by noise level, and how often the agent tells the caller about the noise

Latency

Component

Median response time and the slowest 5% of responses on live calls

Patient acceptance

Task

Completion rate, early requests for a human, sentiment across the call

How Do You Make Sure a Voice Agent Is Safe for Patients?

Every voice agent needs protection against hate speech, discrimination, and abusive language, in both directions. Standard content filters handle these risks, and they belong in every deployment.

Healthcare adds four risks that content filters do not cover:

  • PHI leakage: reading a medication list to a caller before identity verification, or to a caller verified as someone else.

  • Tool misuse: a caller talks the agent into canceling another patient's appointment or submitting a refill nobody requested.

  • Medical misinformation: the agent states a wrong dose, a wrong interaction, or a wrong preparation instruction.

  • Medical advice: a patient asks "can I take two tonight?" A refill agent collects the request and routes the clinical question to a nurse.

1. Escalation for urgent symptoms

Clinical leaders should own the list of urgent phrases and symptoms. Test every phrase on the list in scenarios, including mid-sentence and mid-task versions. Measure escalation recall (the share of urgent calls the agent escalated) and time to transfer. When the agent is uncertain, the default action is a transfer to clinical staff.

2. Guardrails at the vulnerable points

Map the workflow step by step and mark each point where the agent touches patient data or takes an action. Identity verification, record lookup, tool calls that write to the EHR, and web search or knowledge retrieval sit at the top of that list. Put a guardrail at each of those points.

Level AI's real-time agent harness hides transactional tools like refill submission until the caller passes authentication. The same harness runs every incoming and outgoing utterance through parallel safety classifiers in 15 to 20 milliseconds. Test those guardrails with the adversarial calls described in safeguarding your virtual agent against malicious attacks, including prompt injection and social engineering attempts.

3. Catch problems during the call

Post-call review finds a wrong dosing statement after the patient has already hung up with it. Checks on the agent's outgoing responses block that statement before it is spoken and replace it with a transfer. Run both. Live checks protect the patient on the call, and post-call review finds the patterns that need a workflow fix.

Track three safety numbers on every release: attack success rate on the adversarial test set, PHI disclosures before verification (the target is zero), and escalation recall on urgent phrases.

How Do You Keep Evaluating the Agent After Go-Live?

Pre-launch testing covers the calls the team predicted. Live traffic brings the rest. Five signals belong on every call from day one.

  • Virtual agent performance: task completion and containment for each workflow.

  • Transfer reasons: why patients ask for a human, split into out of scope, agent failure, patient preference, and urgent escalation. Each reason has a different fix.

  • Sentiment: where in the call sentiment drops, and which step comes right before the drop.

  • Duplicate requests: new encounters created for patients who already had an open request.

  • Repeat callers: patients who call back within 72 hours about the same request. A contained call followed by a callback is a failed call, one of the failures hiding inside AI agent containment numbers.

Scoring every call by hand is impossible at contact center volume. Auditing virtual agents with VA Pulse scores each virtual agent conversation on agent decisioning, performance and safety, and conversation quality, and marks steps that did not apply to a call as not applicable so the score stays fair.

1. What the refill pilot found after launch

Live call data exposed two problems that pre-launch testing had not.

First, the agent was running full identity verification on calls it could not handle, such as prior authorizations, imaging results, and fax requests. Patients sat through verification only to be transferred. The team moved intent detection to the start of the call, and out-of-scope calls dropped from 90 seconds to under 30 before transfer.

Second, pharmacies calling about refills were mixed into the patient refill volume, which pulled the deflection rate down. Custom analytics dashboards let the contact center leader filter out pharmacy calls and see the deflection rate for the in-scope patient workflows on its own.

2. Measure adoption through pilot usage

Patient adoption shows up in usage data. The refill pilot processed 165 interactions on its first full day of deployment, and the provider saved 983 hours of total conversation time. Track how many patients complete the workflow, how many opt out early, and how those numbers move week over week. Usage data answers the acceptance question more reliably than a pre-launch survey.

Will the Agent Work With Your EHR and Existing Workflows?

EHR integration sets the limit on what the agent does for patients. Three jobs depend on it:

  • Patient verification against the record, before any PHI is shared.

  • Scheduling based on provider rules, such as visit types, slot lengths, and new versus established patient policies.

  • Refill processing against the active medication list and the pharmacy on file.

Evaluate each job with your own provider rules and your own data. Then test the failure cases: an EHR timeout, a patient not found, and a date of birth that does not match the record.

In the primary care pilot, the agent verified patients by phone number and date of birth. It then pulled active medications and pharmacy details from eClinicalWorks and updated Salesforce with the call documentation, with Five9 as the phone system. Every one of those connections needed its own test cases before launch.

1. Developer effort to build and maintain

Ask each vendor how many workflow changes need an engineer. A dialog tree coded branch by branch turns every new provider rule into a development ticket. Level AI configures workflows with deterministic flows for regulated steps like identity verification and controlled open dialogue for the rest of the call. Level AI integrations include 70+ connectors across telephony, EHR and EMR platforms, and CRMs, so the agent connects to the systems already in place.

What Should Each Leader Check?

Each leader on the buying team owns a different part of the evaluation. Use this list to divide the work.

1. IT and digital leaders

  • Architecture: where transcription, reasoning, and guardrails run, and where the audit trail lives.

  • Latency: median and slowest-case response times on live calls, measured on your telephony.

  • Integration: read and write access to the EHR and CRM, and behavior when either system fails.

  • Data safety: PHI handling, encryption, data retention, and HIPAA compliance.

2. Contact center and operations leaders

  • Task completion by workflow, with pharmacy and out-of-scope calls separated.

  • Transfer reasons, and how they shift after each release.

  • Duplicate encounters and repeat callers within 72 hours.

  • Sentiment across the call, tied to the step where it drops.

3. Product leaders

  • Scenario coverage across accents, ages, noise levels, and multi-part calls.

  • A metric structure with task, rubric, and component levels that stays consistent release to release.

  • Regression testing on the full scenario library before every change goes live.

4. Clinical leaders

  • The escalation list, reviewed and signed off by clinical staff.

  • Hard limits on medical advice, with a transfer path for every clinical question.

  • Misinformation checks on dosing, interactions, and preparation instructions.

  • A regular review of escalated calls to confirm the right patients reached the right staff.

How Level AI Keeps Healthcare Voice Agents Accountable After Go-Live

The quality of evaluation decides whether a voice agent improves patient access or adds work for clinical staff. A healthcare contact center needs scenarios drawn from its own patients, metrics that trace a failed call to the step that caused it, and guardrails placed at the points where the agent touches PHI or takes action.

Level AI runs the AI Virtual Agent and the analytics that score it on the same conversation data. VA Pulse scores every virtual agent call, dashboards separate in-scope workflows from the noise, and transfer reasons and sentiment feed the next release. The primary care pilot used that loop to reach 45% refill automation and cut out-of-scope transfers from 90 seconds to under 30.

See how Level AI evaluates healthcare voice agents on every live call.

Level AI scores each virtual agent conversation for task completion, safety, and patient sentiment, then shows where calls break down. Walk through the refill pilot workflow with our team and map it to your own call types.

See how Level AI evaluates healthcare voice agents on every live call.

Level AI scores each virtual agent conversation for task completion, safety, and patient sentiment, then shows where calls break down. Walk through the refill pilot workflow with our team and map it to your own call types.

1. Can the voice agent handle my patients when they ask several questions in one call?

Test multi-part calls directly. Build scenarios where a patient asks for two refills, a lab result, and an appointment change in one call, then measure task completion for each request separately. A call counts as successful only when every in-scope request is complete and every out-of-scope request reaches the right team.

2. How will I know if the agent catches an urgent medical issue and escalates it?

Start with an escalation list your clinical team signs off on, and test every phrase on it in scenarios, including mid-sentence versions. Guardrails on the agent's workflow stop the current task and transfer the patient with context attached. After launch, review escalated calls and track escalation recall, so you see both the urgent calls the agent caught and the ones it missed. The risks of AI in healthcare make this review a standing task for clinical leaders.

3. Will the agent create duplicate requests or telephone encounters for my team?

The agent should check for an existing record before it creates a new one. In the primary care pilot, duplicate documentation appeared when Salesforce did not surface existing telephone encounters, so the lookup step needs its own test cases. After launch, track new encounters created for patients who already had an open request.

4. What happens when my patients call from a noisy place?

Test noisy audio in scenarios at several noise levels, and track word error rate for each level, with medication names tracked separately. When noise is causing errors, the agent should tell the caller so the patient can move somewhere quieter. Watch sentiment and transfer reasons on noisy calls after launch to see where the agent still struggles.

5. Will my patients actually be comfortable talking to an AI?

Tell patients at the start of the call that they are speaking with an AI agent, and give them a clear way to reach a human at any point. Then measure adoption in a pilot: completion rates, early requests for a human, and call center sentiment analysis across every call. The refill pilot processed 165 interactions on its first full day, which gave the team usage data within 24 hours.

table of contents

SHARE THIS POST

Subscribe to Ctrl+CX

Hear insights directly from Rob Dwyer, Level AI's CX Executive in Residence