Back

NigBench: A multilingual point-of-care medical query benchmarking study of large language models in Nigeria

Olatunji, T.; Aka, C.; Okocha, C.; Ayodele, E.; Orisakwe, J.; Adekunle, T.; Sanni, M.; Abiola, A.; Abdullahi, T.; Owopetu, O.; Afolaranmi, T.; Yougha, P. S.; Emmanuel-Fabula, M.; Menon, V.; Denniston, A.; Liu, X.; Williams, G.; Mateen, B. A.

2026-07-10 health informatics
10.64898/2026.07.05.26356776 medRxiv
Show abstract

In this study, we introduce a novel benchmark comprising over 9,000 real-world, point-of-care, multilingual, and multimodal clinical question-answer pairs sourced from frontline health workers in Nigeria. Using the dataset, we compare local general practitioners to multiple leading open and closed LLMs. Our results reveal several critical insights into the suitability of LLMs as clinical decision support systems in low-resource contexts. The results confirm that performance varies widely by language and input modality (e.g., text vs speech): while models perform best on English text inputs, their accuracy drops significantly for local-language speech. Critically, it is possible to achieve substantial performance gains by transcribing and translating other languages into English before prompting an LLM-- an important insight for non-anglophone product developers. Finally, this benchmark highlights key limitations of SLMs in supporting frontline healthcare in low-resource settings and provides a clear opportunity to track improvements as novel solutions are developed.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.