NigBench: A multilingual point-of-care medical query benchmarking study of large language models in Nigeria
Olatunji, T.; Aka, C.; Okocha, C.; Ayodele, E.; Orisakwe, J.; Adekunle, T.; Sanni, M.; Abiola, A.; Abdullahi, T.; Owopetu, O.; Afolaranmi, T.; Yougha, P. S.; Emmanuel-Fabula, M.; Menon, V.; Denniston, A.; Liu, X.; Williams, G.; Mateen, B. A.
Show abstract
In this study, we introduce a novel benchmark comprising over 9,000 real-world, point-of-care, multilingual, and multimodal clinical question-answer pairs sourced from frontline health workers in Nigeria. Using the dataset, we compare local general practitioners to multiple leading open and closed LLMs. Our results reveal several critical insights into the suitability of LLMs as clinical decision support systems in low-resource contexts. The results confirm that performance varies widely by language and input modality (e.g., text vs speech): while models perform best on English text inputs, their accuracy drops significantly for local-language speech. Critically, it is possible to achieve substantial performance gains by transcribing and translating other languages into English before prompting an LLM-- an important insight for non-anglophone product developers. Finally, this benchmark highlights key limitations of SLMs in supporting frontline healthcare in low-resource settings and provides a clear opportunity to track improvements as novel solutions are developed.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Natural language processing to evaluate texting conversations between patients and healthcare providers during COVID-19 Home-Based Care in Rwanda at scale 95%
- Evaluating Anti-LGBTQIA+ Medical Bias in Large Language Models 95%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 95%
Similar papers in this journal
- Using a Multilingual AI Care Agent to Reduce Disparities in Colorectal Cancer Screening: Higher FIT Test Adoption Among Spanish-Speaking Patients 93%
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 93%
- Understanding how the design and implementation of Online Consultations influence primary care outcomes: Systematic review of evidence with recommendations for designers, providers, and researchers 93%
Similar papers in this journal
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 93%
- AI-Generated Clinical Summaries: Errors and Susceptibility to Speech and Speaker Variability 93%
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 92%
Similar papers in this journal
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 94%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 93%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.