Comparison of Large Language Models' Performance on Neurosurgical Board Examination Questions
Andrade, N. S.; Donty, S. V.
Show abstract
BackgroundMultiple-choice board examinations are a primary objective measure of competency in medicine. Large language models (LLMs) have demonstrated rapid improvements in performance on medical board examinations in the past two years. We evaluated five leading LLMs on neurosurgical board exam questions. MethodsWe evaluated five LLMs (OpenAI o1, OpenEvidence, Claude 3.5 Sonnet, Gemini 2.0, and xAI Grok2) on 500 multiple-choice questions from the Self-Assessment in Neurological Surgery (SANS) American Board of Neurological Surgery (ABNS) Primary Board Examination Review. Performance was analyzed across 12 subspecialty categories and compared to established passing thresholds. ResultsAll models exceeded the threshold for passing, with OpenAI o1 achieving the highest accuracy (87.6%), followed by OpenEvidence (84.2%), Claude 3.5 Sonnet (83.2%), Gemini 2.0 (81.0%) and xAI Grok2 (79.0%). Performance was strongest in Other General (97.4%) and Peripheral Nerve (97.1%) categories, while Neuroradiology showed the lowest accuracy (57.4%) across all models. ConclusionsState of the art LLMs continue to improve, and all models demonstrated strong performance on neurosurgical board examination questions. Medical image analysis continues to be a limitation of current LLMs. The current level of LLM performance challenges the relevance of written board examinations in trainee evaluation and suggests that LLMs are ready for implementation in clinical medicine and medical education.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Large language models for generating medical examinations: systematic review 95%
- Performance of ChatGPT on Chinese National Medical Licensing Examinations: A Five-Year Examination Evaluation Study for Physicians, Pharmacists and Nurses 94%
- Medical students' perceptions towards artificial intelligence in education and practice: A multinational, multicenter cross-sectional study 92%
Similar papers in this journal
Similar papers in this journal
- tbiExtractor: A framework for Extracting Traumatic Brain Injury Common Data Elements from Radiology Reports 92%
- ai-corona : Radiologist-Assistant Deep Learning Framework for COVID-19 Diagnosis in Chest CT Scans 92%
- Enhancing Semantic Segmentation in Chest X-Ray Images through Image Preprocessing: ps-KDE for Pixel-wise Substitution by Kernel Density Estimation 91%
Similar papers in this journal
- Performance of ChatGPT and GPT-4 on Neurosurgery Written Board Examinations 96%
- Performance of ChatGPT, GPT-4, and Google Bard on a Neurosurgery Oral Boards Preparation Question Bank 94%
- Low and Borderline Ankle Brachial Index is Associated with Intracranial Aneurysms – a Retrospective Cohort Study 83%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.