Multimodal Large Language Model Passes Specialty Board Examination and Surpasses Human Test-Taker Scores: A Comparative Analysis Examining the Stepwise Impact of Model Prompting Strategies on Performance
Samaan, J. S.; Margolis, S.; Srinivasan, N.; Srinivasan, A.; Yeo, Y. H.; Anand, R.; Samaan, F. S.; Mirocha, J.; Safavi-Naini, S. A. A.; El Kurdi, B.; Soroush, A.; Watson, R.; Gaddam, S.; Elmore, J. G.; Spiegel, B. M. R.; Tatonetti, N. P.
Show abstract
BackgroundLarge language models (LLMs) have shown promise in answering medical licensing examination-style questions. However, there is limited research on the performance of multimodal LLMs on subspecialty medical examinations. Our study benchmarks the performance of multimodal LLMs enhanced by model prompting strategies on gastroenterology subspeciality examination-style questions and examines how these prompting strategies incrementally improve overall performance. MethodsWe used the 2022 American College of Gastroenterology (ACG) self-assessment examination (N=300). This test is typically completed by gastroenterology fellows and established gastroenterologists preparing for the gastroenterology subspeciality board examination. We employed a sequential implementation of model prompting strategies: prompt engineering, retrieval augmented generation (RAG), five-shot learning, and an LLM-powered answer validation revision model (AVRM). GPT-4 and Gemini Pro were tested. ResultsImplementing all prompting strategies improved the overall score of GPT-4 from 60.3% to 80.7% and Gemini Pros from 48.0% to 54.3%. GPT-4s score surpassed the 70% passing threshold and 75% average human test-taker scores unlike Gemini Pro. Stratification of questions by difficulty showed the accuracy of both LLMs mirrored that of human examinees, demonstrating higher accuracy as human test-taker accuracy increased. The addition of the AVRM to prompt, RAG and 5-shot increased GPT-4s accuracy by 4.4%. The incremental addition of model prompting strategies improved accuracy for both non-image (57.2% to 80.4%) and image-based (63.0% to 80.9%) questions for GPT-4, but not Gemini Pro. ConclusionsOur results underscore the value of model prompting strategies in improving LLM performance on subspecialty-level licensing exam questions. We also present a novel implementation of an LLM-powered reviewer model in the context of subspecialty medicine which further improved model performance when combined with other prompting strategies. Our findings highlight the potential future role of multimodal LLMs, particularly with the implementation of multiple model prompting strategies, as clinical decision support systems in subspecialty care for healthcare providers.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 91%
- Bridging the Literacy Gap for Surgical Consents: An AI-Human Expert Collaborative Approach 91%
- International Electronic Health Record-Derived COVID-19 Clinical Course Profiles: The 4CE Consortium 91%
Similar papers in this journal
Similar papers in this journal
- Assessing ChatGPT’s Mastery of Bloom’s Taxonomy using psychosomatic medicine exam questions 92%
- Using a Multilingual AI Care Agent to Reduce Disparities in Colorectal Cancer Screening: Higher FIT Test Adoption Among Spanish-Speaking Patients 91%
- One LLM is not Enough: Harnessing the Power of Ensemble Learning for Medical Question Answering 90%
Similar papers in this journal
- irAE-GPT: Leveraging large language models to identify immune-related adverse events in electronic health records and clinical trial datasets 91%
- Enhancing Early Detection of Cognitive Decline in the Elderly through Ensemble of NLP Techniques: A Comparative Study Utilizing Large Language Models in Clinical Notes 90%
- Consistent Performance of GPT-4o in Rare Disease Diagnosis Across Nine Languages and 4967 Cases 90%
Similar papers in this journal
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 93%
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 92%
- Usability of a Machine-Learning Clinical Order Recommender System Interface for Clinical Decision Support and Physician Workflow 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.