A 'Silent Trial' Assessing the Accuracy of Large Language Models for Assisting Community Health Workers in Low-Resource Settings
Shimelash, N.; Rutunda, S.; Menon, V.; Emmanual-Fabula, M.; Uwimbabazi, A.; Rugege, C.; Nshimiyimana, C.; Rwema, I.; Kandekwe, M.; Berhe, D. F. D.; Wong, R.; Remera, E.; Hezagira, E.; Gill, J.; Archer, L.; Riley, R. D.; Denniston, A. K.; Liu, X.; Mateen, B.
Show abstract
Community health workers (CHWs) in low-resource settings deliver variable-quality care. This study used OpenAIs o3 and Googles Gemini Flash 2.5 to evaluate whether large language models (LLMs) listening to CHW-patient interactions could generate accurate referral decisions. Across 150 participating Rwandan CHWs, 429 encounters were recorded (in Kinyarwanda) and then processed by LLMs. CHWs demonstrated high referral accuracy (97.9% [95% CI: 96.1%-98.9%]), and OpenAIs o3 performed similarly to CHWs while Gemini 2.5-Flash showed low accuracy (47.3% [95% CI: 42.6%-52.1%]). Assessment of LLM-generated differential diagnoses and management plan quality showed superior performance from o3 compared with Gemini, though both models missed important conditions. In conclusion, the choice of LLM appears to be a critical design decision. Moreover, the high baseline performance of Rwandan CHWs suggests that LLMs are likely to have a limited impact in the current context but could be useful in less well-established CHW programmes. Trial Registration: PACTR202504601308784.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Self-tests for COVID-19: what is the evidence? A living systematic review and meta-analysis (2020-2023) 95%
- Waiting times, patient flow, and occupancy density in South African primary health care clinics: implications for infection prevention and control 94%
- Improving the quality of in-patient neonatal routine data as a pre-requisite for monitoring and improving quality of care at scale: A multi-site retrospective cohort study in Kenyan hospitals 94%
Similar papers in this journal
- Identifying clinical skill gaps of healthcare workers using a digital clinical decision support algorithm during outpatient pediatric consultations in primary health centers in Rwanda 96%
- “It reminds me and motivates me” : Human-centered design and implementation of an interactive, SMS-based digital intervention to improve early retention on antiretroviral therapy: usability and acceptability among new initiates in a high-volume, public clinic in Malawi 94%
- Wellbeing Impact Study of High-Speed 2 (WISH2): Protocol for a mixed-methods examination of the impact of major transport infrastructure development on mental health and wellbeing 94%
Similar papers in this journal
- Understanding how the design and implementation of Online Consultations influence primary care outcomes: Systematic review of evidence with recommendations for designers, providers, and researchers 95%
- Using a Multilingual AI Care Agent to Reduce Disparities in Colorectal Cancer Screening: Higher FIT Test Adoption Among Spanish-Speaking Patients 94%
- Improving Patient Engagement in Phase 2 Clinical Trials with a Trial-specific Patient Decision Aid (tPDA): A Development and Usability Study 94%
Similar papers in this journal
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 95%
- The performance of national COVID-19 ‘Symptom Checkers’: A comparative case simulation study 94%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.