Evaluation of Large Language Models in Medical Examinations:A Scoping Review Protocol
Wang, W.; Wang, B.; Zhu, Y.; Wang, Z.; peng, S.
Show abstract
IntroductionLarge language models (LLMs) demonstrate human-level performance in three key domains: linguistic understanding, knowledge-based reasoning, and complex problem-solving. These characteristics make LLMs valuable tools for medical education. Standardized medical examinations evaluate clinical competencies in trainees. These examinations allow rigorous verification of LLMs accuracy and reliability in medical contexts. Current methods use standardized examinations to test LLMs clinical reasoning abilities. Significant performance variations emerge across different clinical scenarios. No comprehensive reviews have compared different LLM versions in medical examinations. Most studies focus on individual models, lacking comparative analyses of multiple LLM versions. Current approaches struggle to keep pace with evolving research needs. This study synthesizes extant research on LLMs in medical examinations, by analyzing the current challenges and limitations, offers guidance for future investigations. Methods and analysisThe protocol was designed following the JBI Manual for Evidence Synthesis guidelines. We established explicit inclusion/exclusion criteria and search strategies. Systematic searches were performed in PubMed and Web of Science Core Collection databases. The methodology details literature screening, data extraction, analysis frameworks, and process mapping. This approach ensures methodological rigor throughout the research process. Ethics and disseminationThis protocol outlines a scoping review methodology. The study involves systematic synthesis and analysis of published literature. It does not include human/animal experimentation or sensitive data collection. Ethical approval is not required for this literature-based study. Strengths and limitations of this studyO_LIThis scoping review programme strictly adheres to the standardized guidelines for the implementation of scoping reviews. Includes the JBI Manual for Evidence Synthesis and the Preferred Reporting Items for Systematic Reviews and Scoping Reviews Extended Meta-Analysis (PRISMA-ScR) guideline. C_LIO_LIThe search strategy included two databases:PubMed, Web of Science Core Collection. C_LIO_LIThis scoping review will bridge the knowledge gap of LLMs across medical examinations due to recent rapid technological advances. C_LIO_LIBy the nature of the scoping review, failure to critically evaluate identified sources of evidence. C_LI The results of the scoping review will serve as a basis for identifying directions for further research on LLMs in the field of medical examinations.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Large language models for generating medical examinations: systematic review 97%
- Medical students' perceptions towards artificial intelligence in education and practice: A multinational, multicenter cross-sectional study 96%
- Performance of ChatGPT on Chinese National Medical Licensing Examinations: A Five-Year Examination Evaluation Study for Physicians, Pharmacists and Nurses 95%
Similar papers in this journal
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 95%
- Impact of electronic medical records on healthcare delivery in Nigeria: A Review 94%
- Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study 94%
Similar papers in this journal
- Evaluation of Self-Directed Learning Activities at King Abdulaziz University: A Qualitative Study of Faculty Perceptions 94%
- Impact of mHealth interventions on antenatal and postnatal care utilization in low and middle-income countries: A Systematic Review and Meta-Analysis 92%
- Validation of the patient reported outcome measures tool “Catquest” in Odia language 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.