How Good Are Large Language Models at Supporting Frontline Healthcare Workers in Low-Resource Settings: A Benchmarking Study & Dataset
Rutunda, S.; Williams, G.; Kabanda, K.; Nkuruniz, F.; Uwiduhaye, S.; Rugegamanzi, E.; Nshimiyimana, C.; Menon, V.; Emmanuel-Fabula, M.; Liu, X.; Hezagira, E.; Mateen, B. A.
Show abstract
Large language models (LLMs) have demonstrated strong performance in medical contexts; however, existing benchmarks often fail to reflect the real-world complexity of low-resource health systems accurately. This study developed a dataset of 5,609 clinical questions contributed by 101 community health workers (CHWs) across four Rwandan districts and compared responses generated by five large language models (LLMs) (Gemini-2, GPT-4o, o3 mini, Deepseek R1, and Meditron-70B) with those from local clinicians. A subset of 524 question-answer pairs was evaluated using a rubric of 11 expert-rated metrics, scored on a five-point Likert scale. Gemini-2 and GPT-4o were the best performers (achieving mean scores of 4.49 and 4.48 out of 5, respectively, across all 11 metrics). All LLMs significantly outperformed local clinicians (ps < 0.001) across all metrics, with Gemini-2, for example, surpassing local GPs by an average of 0.83 points on every metric (range: 0.38 - 1.10). While performance degraded slightly when LLMs communicated in Kinyarwanda, the LLMs remained superior to clinicians and were over 500 times cheaper per response. These findings support the potential of LLMs to strengthen frontline care quality in low-resource, multilingual health systems.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Collecting mortality data via mobile phone surveys: a non-inferiority randomized trial in Malawi 92%
- How does policy modelling work in practice? A global analysis on the use of epidemiological modelling in health crises 92%
- Waiting times, patient flow, and occupancy density in South African primary health care clinics: implications for infection prevention and control 92%
Similar papers in this journal
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 93%
- Identifying clinical skill gaps of healthcare workers using a digital clinical decision support algorithm during outpatient pediatric consultations in primary health centers in Rwanda 93%
- Wellbeing Impact Study of High-Speed 2 (WISH2): Protocol for a mixed-methods examination of the impact of major transport infrastructure development on mental health and wellbeing 93%
Similar papers in this journal
- Ethnicity and COVID-19 outcomes among healthcare workers in the United Kingdom: UK-REACH ethico-legal research, qualitative research on healthcare workers’ experiences, and stakeholder engagement protocol 93%
- What is the suitability of clinical vignettes in benchmarking the performance of online symptom checkers? An audit study 93%
- Implementation Science Protocol for a participatory, theory-informed implementation research programme in the context of health system strengthening in sub-Saharan Africa (ASSET-ImplementER) 93%
Similar papers in this journal
Similar papers in this journal
- Using a Multilingual AI Care Agent to Reduce Disparities in Colorectal Cancer Screening: Higher FIT Test Adoption Among Spanish-Speaking Patients 92%
- Understanding how the design and implementation of Online Consultations influence primary care outcomes: Systematic review of evidence with recommendations for designers, providers, and researchers 92%
- Virtual health care for community management of patients with COVID-19. 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.