Multilingual Evaluation of a Large Language Model-Based Primary Care Chatbot
Chen, P.-L.; Rao, A. A.; Pugh, S. F.; Johnson, K. B.
Show abstract
Pre-visit planning has the potential to reduce EHR documentation burden while improving workflow efficiency, care quality, and patient-provider engagement. Large language model (LLM) chatbots show promise for supporting this task, but while their English-centric development suggests a potential for disparity, the extent to which these concerns translate into performance degradation in multilingual clinical settings remains unclear. In this mixed-methods study, we systematically evaluate the multilingual capabilities of PCP-Bot, an English-developed LLM-based (GPT-4o) clinical chatbot that collects patient concerns and generates structured, physician-ready summaries ([~]200 words) under structured output constraints. We enrolled 31 bilingual individuals (11 Mandarin, 10 Spanish, 10 Hindi) to role-play as patients to evaluate the PCP-Bot, interacting with it across five synthetic clinical cases in both English and a second language. Participants completed a structured survey comprising baseline language proficiency screening, standardized interactions with PCP-Bot in each language, and post-interaction evaluations. Case order was randomized, with each scenario completed first in English and subsequently in the participants second language. All summaries were generated in English, regardless of the interaction language. Our results show that Hindi achieved usability and conversation quality parity with English across all measured dimensions. Mandarin achieved usability parity but showed a significant conversation quality gap relative to English. Spanish demonstrated significant deficits in both conversation quality and summary quality. Trust and workload remained consistent across languages. Qualitatively, participants found PCP-Bot natural, smooth, and accurate overall, but noted repetition, transcription errors, missed follow-ups, and more frequent usability issues in non-English interactions. Overall, our findings demonstrate that LLM translation capabilities can enable effective deployment beyond English following appropriate performance validation.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Improving Patient Engagement in Phase 2 Clinical Trials with a Trial-specific Patient Decision Aid (tPDA): A Development and Usability Study 96%
- Using a Multilingual AI Care Agent to Reduce Disparities in Colorectal Cancer Screening: Higher FIT Test Adoption Among Spanish-Speaking Patients 94%
- Understanding how the design and implementation of Online Consultations influence primary care outcomes: Systematic review of evidence with recommendations for designers, providers, and researchers 94%
Similar papers in this journal
Similar papers in this journal
- A digital self-care intervention for Ugandan patients with heart failure and their clinicians: User-centred design and usability study 94%
- Validating a Clinical Decision Support System for Palliative Care using healthcare professionals’ insights 94%
- How suitable are clinical vignettes for the evaluation of symptom checker apps? A test theoretical perspective 93%
Similar papers in this journal
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 95%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 94%
- Utilization of Generative AI-drafted Responses for Managing Patient-Provider Communication 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.