Large language models for self-administered conversational vignette assessment of provider competencies: A pilot and validation study in Vietnam with automated LLM-powered transcript classification
Daniels, B.; Zhang, W.; Nguyen, H.; Duong, D.
Show abstract
We developed and validated a self-administered clinical vignette platform powered by a large language model (LLM), deployed through a SurveyCTO web survey, to measure primary health care provider competencies in Vietnam. In a pilot focus group, nine physicians rated LLM-simulated patient interactions as realistic (mean 3.78/5) and user-friendly. In the validation phase, 22 providers completed 132 vignette interactions across ten clinical scenarios in Vietnamese. Essential diagnostic checklist scores (human-coded from translated transcripts) correlated with expert clinician evaluations (Pearsons{rho} = 0.55-0.60). LLM-automated coding of checklist items from translated English transcripts correlated reasonably with human coding ({rho} = 0.53), and coding directly from Vietnamese transcripts performed comparably ({rho} = 0.51), suggesting that a separate translation step may not be necessary. The total cost of 132 chatbot interactions was under USD 2. LLM-driven conversational vignettes represent a low-cost and scalable method for assessing provider competencies in respondents local language, eliminating the need for extensive enumeration staffs while preserving the open-ended format critical to vignette validity, and additionally introducing flexible feature extraction from transcripts using grading rubrics. The platform is open-source and designed for replication in other health system contexts. Author summaryMeasuring the clinical skills of healthcare providers is essential for improving the quality of care, but current survey methods are expensive and require trained enumerators to travel to health facilities in person. We developed a new approach that uses large language models (LLMs) - the technology behind tools like ChatGPT and Claude - to simulate patients in realistic clinical conversations that healthcare providers can complete on their phones or laptops over the Internet in their own language. In Vietnam, we tested this tool with 31 physicians across ten clinical scenarios. Providers found the simulated patient conversations realistic and easy to use. We also tested whether LLMs could automatically score the conversations, which showed reasonable agreement with human scoring, and performed nearly as well when scoring directly from Vietnamese, without requiring a separate translation step. When we compared these results from our tool against holistic expert physician ratings of the same conversations, the scores agreed well, suggesting that automatic transcript grading based on rubrics produces meaningful measures of clinical skill. This tool costs less than two US dollars for over a hundred consultations and required no in-person surveyors, making it potentially transformative for routine, large-scale monitoring of healthcare quality in resource-limited settings. The platform and code are openly available for adaptation.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Using a Multilingual AI Care Agent to Reduce Disparities in Colorectal Cancer Screening: Higher FIT Test Adoption Among Spanish-Speaking Patients 95%
- Understanding how the design and implementation of Online Consultations influence primary care outcomes: Systematic review of evidence with recommendations for designers, providers, and researchers 95%
- Improving Patient Engagement in Phase 2 Clinical Trials with a Trial-specific Patient Decision Aid (tPDA): A Development and Usability Study 94%
Similar papers in this journal
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 95%
- Evaluating user experience with immersive technology in simulation-based education: a modified Delphi study with qualitative analysis 94%
- Identifying clinical skill gaps of healthcare workers using a digital clinical decision support algorithm during outpatient pediatric consultations in primary health centers in Rwanda 94%
Similar papers in this journal
- What is the suitability of clinical vignettes in benchmarking the performance of online symptom checkers? An audit study 93%
- Physician experiences of electronic health records interoperability and its practical impact on care delivery in the English NHS: A cross-sectional survey study 93%
- Ethnicity and COVID-19 outcomes among healthcare workers in the United Kingdom: UK-REACH ethico-legal research, qualitative research on healthcare workers’ experiences, and stakeholder engagement protocol 93%
Similar papers in this journal
- Self-tests for COVID-19: what is the evidence? A living systematic review and meta-analysis (2020-2023) 94%
- Improving patient-centred counselling skills among lay healthcare workers in South Africa using the Thusa-Thuso motivational interviewing training and support program 93%
- Comparing in-person, blended and virtual training interventions; a real-world evaluation of HIV capacity building programs in 16 countries in sub-Saharan Africa 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.