Back

Performance of ChatGPT on free-response, clinical reasoning exams

Strong, E.; DiGiammarino, A.; Weng, Y.; Basaviah, P.; Hosamani, P.; Kumar, A.; Nevins, A.; Kugler, J.; Hom, J.; Chen, J.

2023-03-29 medical education
10.1101/2023.03.24.23287731 medRxiv
Show abstract

ImportanceStudies show that ChatGPT, a general purpose large language model chatbot, could pass the multiple-choice US Medical Licensing Exams, but the models performance on open-ended clinical reasoning is unknown. ObjectiveTo determine if ChatGPT is capable of consistently meeting the passing threshold on free-response, case-based clinical reasoning assessments. DesignFourteen multi-part cases were selected from clinical reasoning exams administered to pre-clerkship medical students between 2019 and 2022. For each case, the questions were run through ChatGPT twice and responses were recorded. Two clinician educators independently graded each run according to a standardized grading rubric. To further assess the degree of variation in ChatGPTs performance, we repeated the analysis on a single high-complexity case 20 times. SettingA single US medical school ParticipantsChatGPT Main Outcomes and MeasuresPassing rate of ChatGPTs scored responses and the range in model performance across multiple run throughs of a single case. Results12 out of the 28 ChatGPT exam responses achieved a passing score (43%) with a mean score of 69% (95% CI: 65% to 73%) compared to the established passing threshold of 70%. When given the same case 20 separate times, ChatGPTs performance on that case varied with scores ranging from 56% to 81%. Conclusions and RelevanceChatGPTs ability to achieve a passing performance in nearly half of the cases analyzed demonstrates the need to revise clinical reasoning assessments and incorporate artificial intelligence (AI)-related topics into medical curricula and practice.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.