Probing the Surgical Competence of LLMs: A global health study leveraging AfriMedQA benchmarks
Olatunji, T.; Omofoye, F.; Aka, E. C.; Itzikowitz, G.; Macaulay, D.; Adaramola, O. G.; Adewale, B. A.; Asuzu, C.; Popoola, S.; Kinara, W.; Ayodele, E.; Yinusa, I.; Adekunle, O.; Sanni, M.; Okocha, C.; Abdullahi, T.; Owodunni, A.; Nimo, C.; Asiedu, M. N.; Mateen, B.; Weintraub, R.
Show abstract
Global surgical care faces a severe workforce shortage, with more than 1.2 million additional specialists needed by 2030, particularly in low- and middle-income countries (LMICs). Large language models (LLMs) have demonstrated impressive medical reasoning on standardized exams, but their safety, reliability, and specialty-specific performance--especially in procedural fields such as surgery--remain uncertain. Here we evaluate over 40 state-of-the-art LLMs on 3,900 expert-authored multiple-choice questions across 32 medical specialties from the AfriMed-QA benchmark, developed by 20 African medical professors. Top models (o1, GPT-4o, Claude 3.5) achieved mean accuracies exceeding 82%, showing strong diagnostic reasoning, yet consistently underperformed in surgery, pathology, and obstetrics compared with medical disciplines. Error analyses revealed frequent procedural reasoning failures, omission of local clinical guidelines, and overconfident but incorrect answers. Smaller or biomedical models exhibited higher hallucination and formatting error rates, while prompting strategies had inconsistent benefits. These results highlight the uneven readiness of LLMs for specialty-specific decision support and underscore the need for locally grounded evaluation frameworks, improved instruction tuning, and rigorous real-world validation to ensure the safe and equitable deployment of AI-assisted clinical tools in LMICs.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 95%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 94%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 94%
Similar papers in this journal
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 92%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 90%
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 90%
Similar papers in this journal
Similar papers in this journal
- irAE-GPT: Leveraging large language models to identify immune-related adverse events in electronic health records and clinical trial datasets 93%
- Consistent Performance of GPT-4o in Rare Disease Diagnosis Across Nine Languages and 4967 Cases 93%
- Enhancing Early Detection of Cognitive Decline in the Elderly through Ensemble of NLP Techniques: A Comparative Study Utilizing Large Language Models in Clinical Notes 92%
Similar papers in this journal
- Loon Lens 1.0 Validation: Agentic AI for Title and Abstract Screening in Systematic Literature Reviews 88%
- The roadmap for implementing value based healthcare in European university hospitals - consensus report and recommendations 86%
- Emerging Therapies for COVID-19: the value of information from more clinical trials 85%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.