Back

Evaluating a Medical-Grade Voice AI for Patient and Caregiver Guidance: A Multi-Scenario Nurse Panel Study

Sridhar, S.; Vasantha, R.; Deshpande, S.; Pathak, P.

2025-09-19 health informatics
10.1101/2025.09.18.25336107 medRxiv
Show abstract

Medical-grade conversational AI offers the potential to extend patient support programs (PSPs) in regulated therapeutic areas, but its safety and reliability must be rigorously evaluated. We conducted a large-scale, nurse-led assessment of a voice-based AI system across 30 patient and caregiver scenarios spanning diabetes, oncology, neurology, cardiometabolic disease, and rare disorders. Nearly 1,000 U.S.-licensed nurses role-played patients or caregivers in more than 20,000 interactions, scoring the AI across five domains: clinical accuracy, empathy, communication clarity, appropriateness of advice, and compliance with evidence. The system achieved over 97% top ratings across all domains, with empathy noted most strongly in oncology and rare disease caregiving contexts, and clarity reflecting minor opportunities in pacing and call-closing behavior. Qualitative feedback emphasized tone, personalization, and regulatory compliance as consistent strengths. These findings demonstrate that voice-based AI can safely and effectively support patient and caregiver interactions, suggesting readiness for scaled deployment in pharma-led PSPs to expand after-hours coverage, provide consistent patient engagement, and generate feedback loops to inform both AI and human nurse training.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.