Benchmarking Large Language Models and Clinicians Using Locally Generated Primary Healthcare Vignettes in Kenya
Mwaniki, P.; Musau, W.; Isaaka, L.; Wanyama, C.; Menon, V.; Denniston, A. K.; Liu, X.; Emmanual-Fabula, M.; Williams, G.; Mateen, B. A.; Agweyu, A.
Show abstract
BackgroundLarge language models (LLMs) show promise on healthcare tasks, yet most evaluations emphasize multiple-choice accuracy rather than open-ended reasoning. Evidence from low-resource settings remains limited. MethodsWe benchmarked five LLMs (GPT-4.1, Gemini-2.5-Flash, DeepSeek-R1, MedGemma, and o3) against Kenyan clinicians, using a randomly subsampled dataset of 507 vignettes (from a larger pool of 5,107 clinical scenarios) spanning 12 nursing competency categories. Blinded physician panels rated responses using a 5- point Likert scale on an 11-domain rubric covering accuracy, safety, contextual appropriateness, and communication. We summarized mean scores and used Bayesian ordinal logistic regression to estimate probabilities of high-quality ratings ([≥]4) and to perform pairwise comparisons between LLMs and clinicians. FindingsClinician mean ratings were lower than those for LLMs in 9/11 domains: 2.86 vs 4.25-4.72 (guideline alignment), 2.76 vs 4.25-4.73 (expert knowledge), 2.96 vs 4.30-4.73 (logical coherence), and 2.58 vs 4.16-4.68 (low omission of critical information). On safety-related domains, LLMs received higher ratings: minimal extent of possible harm 3.16 vs 4.29-4.68; low likelihood of harm 3.68 vs 4.54-4.81. Performance was similar for low inclusion of irrelevant content (4.28 vs 4.25-4.35) and for avoidance of demographic bias (4.86 vs 4.91-4.94). In Bayesian models, LLMs had >90% probability of ratings [≥]4 in most domains, whereas clinicians exceeded 90% only for contextual relevance and demographic/socio-economic bias. Pairwise contrasts showed broadly overlapping credible intervals among LLMs, with o3 leading numerically most domains except contextual relevance, demographic/socio-economic bias, and relevance to the question. Generating all LLM responses cost USD 3.86-8.68 per model (USD 0.008-0.017 per vignette), compared with USD 3.35 per clinician-generated vignette. InterpretationLLMs produced responses that were more accurate, safer, and more structured than clinicians in vignette-based tasks. Findings support further evaluation of LLMs as decision support in resource-constrained health systems. Funding StatementThis study was supported by the Gates Foundation [INV-068056]. Research in ContextO_ST_ABSEvidence before this studyC_ST_ABSWe searched PubMed, medRxiv, and arXiv (Jan 1, 2021-Sept 30, 2025) using combinations of terms including "large language model", "LLM", "healthcare", "benchmarking", "clinical decision support", and "low-resource settings". The search returned 28 preprints and only 4 peer-reviewed articles. A study from Rwanda benchmarked five LLMs against clinicians using 524 real-world questions from community health workers; all models outperformed clinicians, including in Kinyarwanda (Rutunda, 2025). In Kenya, a multimodal LLM (POE) outperformed primary care providers on 63 otolaryngology cases (79.4% vs 50.8%) and aligned with specialist recommendations (Lechien, 2025). A cross-country maternal health study evaluated GPT-4, GPT-3.5, a custom GPT-3.5, and Meditron-70b on three questions, with expert reviewers in Brazil, Pakistan, and the USA rating outputs in their native languages. GPT-4 and GPT-3.5 were most accurate, though readability and gender bias were noted (Lima, 2025). AraSum, a lightweight Arabic summarization model, outperformed the Arabic foundation model JAIS-30B on BLEU, ROUGE, and expert ratings of accuracy, comprehensiveness, and clinical utility (Lee, 2025). Additional preprints proposed expert-rated benchmarks for LMIC clinical tasks. Added value of this studyThis study uniquely combines local co-design, real-world clinical scenarios, and structured, expert-based assessment across 11 dimensions of clinical quality. It demonstrates the relative strengths and weaknesses of five widely available LLMs versus frontline clinician performance, offering evidence of systematic clinician gaps in accuracy, guideline adherence, and completeness. Implications of all the available evidenceLLMs show substantial promise as clinical decision support tools in low-resource health systems. Across multiple settings and task types, current models consistently meet or exceed clinician performance in controlled evaluations. However, real-world deployment requires attention to equity, local clinical validation, and thoughtful implementation pathways that mitigate risk and reinforce trust.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Implementing essential diagnostics-learning from essential medicines: A scoping review 93%
- Improving the quality of in-patient neonatal routine data as a pre-requisite for monitoring and improving quality of care at scale: A multi-site retrospective cohort study in Kenyan hospitals 92%
- Architecture of systems affecting disease trajectories in a conflict zone: A community-centered systems inquiry in North Gaza 92%
Similar papers in this journal
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 93%
- Measures of socioeconomic advantage are not independent predictors of support for healthcare AI: subgroup analysis of a national Australian survey 92%
- Cracking the Code: A Scoping Review to Unite Disciplines in Tackling Legal Issues in Health Artificial Intelligence 92%
Similar papers in this journal
- Remote Covid Assessment in Primary Care (RECAP) risk prediction tool: derivation and real-world validation studies 92%
- Automated and partially-automated contact tracing: a rapid systematic review to inform the control of COVID-19 92%
- Automated and semi-automated contact tracing: Protocol for a rapid review of available evidence and current challenges to inform the control of COVID-19 91%
Similar papers in this journal
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 94%
- Essential Indicators of Quality in Primary Care Settings: An Evidence-Based, Structured, Expert Approach 93%
- Development of the Tool for Advancing Practice Performance, a practice-level survey to assess primary care structures and processes 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.