Performance of ChatGPT on free-response, clinical reasoning exams
Strong, E.; DiGiammarino, A.; Weng, Y.; Basaviah, P.; Hosamani, P.; Kumar, A.; Nevins, A.; Kugler, J.; Hom, J.; Chen, J.
Show abstract
ImportanceStudies show that ChatGPT, a general purpose large language model chatbot, could pass the multiple-choice US Medical Licensing Exams, but the models performance on open-ended clinical reasoning is unknown. ObjectiveTo determine if ChatGPT is capable of consistently meeting the passing threshold on free-response, case-based clinical reasoning assessments. DesignFourteen multi-part cases were selected from clinical reasoning exams administered to pre-clerkship medical students between 2019 and 2022. For each case, the questions were run through ChatGPT twice and responses were recorded. Two clinician educators independently graded each run according to a standardized grading rubric. To further assess the degree of variation in ChatGPTs performance, we repeated the analysis on a single high-complexity case 20 times. SettingA single US medical school ParticipantsChatGPT Main Outcomes and MeasuresPassing rate of ChatGPTs scored responses and the range in model performance across multiple run throughs of a single case. Results12 out of the 28 ChatGPT exam responses achieved a passing score (43%) with a mean score of 69% (95% CI: 65% to 73%) compared to the established passing threshold of 70%. When given the same case 20 separate times, ChatGPTs performance on that case varied with scores ranging from 56% to 81%. Conclusions and RelevanceChatGPTs ability to achieve a passing performance in nearly half of the cases analyzed demonstrates the need to revise clinical reasoning assessments and incorporate artificial intelligence (AI)-related topics into medical curricula and practice.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Large language models for generating medical examinations: systematic review 95%
- Performance of ChatGPT on Chinese National Medical Licensing Examinations: A Five-Year Examination Evaluation Study for Physicians, Pharmacists and Nurses 94%
- Evaluation of Statistical Illiteracy in Latin American Clinicians and of the Efficacy of a 10-Hour Course 93%
Similar papers in this journal
Similar papers in this journal
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 94%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 94%
- Utilization of Generative AI-drafted Responses for Managing Patient-Provider Communication 92%
Similar papers in this journal
- Medical Clinical Minds Meet Artificial Intelligence: Italian Physicians' Knowledge, Attitudes, and Concordance between Italian Physicians and AI-Generated Diagnoses. A National Cross-Sectional Study 93%
- Large Language Models in Real-World Clinical Workflows: A Systematic Review of Applications and Implementation 90%
- Development and Validation of a Machine Learning Model Integrated with the Clinical Workflow for Inpatient Discharge Date Prediction 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.