Back

A 'Silent Trial' Assessing the Accuracy of Large Language Models for Assisting Community Health Workers in Low-Resource Settings

Shimelash, N.; Rutunda, S.; Menon, V.; Emmanual-Fabula, M.; Uwimbabazi, A.; Rugege, C.; Nshimiyimana, C.; Rwema, I.; Kandekwe, M.; Berhe, D. F. D.; Wong, R.; Remera, E.; Hezagira, E.; Gill, J.; Archer, L.; Riley, R. D.; Denniston, A. K.; Liu, X.; Mateen, B.

2026-02-17 primary care research
10.64898/2026.02.16.26346409 medRxiv
Show abstract

Community health workers (CHWs) in low-resource settings deliver variable-quality care. This study used OpenAIs o3 and Googles Gemini Flash 2.5 to evaluate whether large language models (LLMs) listening to CHW-patient interactions could generate accurate referral decisions. Across 150 participating Rwandan CHWs, 429 encounters were recorded (in Kinyarwanda) and then processed by LLMs. CHWs demonstrated high referral accuracy (97.9% [95% CI: 96.1%-98.9%]), and OpenAIs o3 performed similarly to CHWs while Gemini 2.5-Flash showed low accuracy (47.3% [95% CI: 42.6%-52.1%]). Assessment of LLM-generated differential diagnoses and management plan quality showed superior performance from o3 compared with Gemini, though both models missed important conditions. In conclusion, the choice of LLM appears to be a critical design decision. Moreover, the high baseline performance of Rwandan CHWs suggests that LLMs are likely to have a limited impact in the current context but could be useful in less well-established CHW programmes. Trial Registration: PACTR202504601308784.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.