Not Such a Long Way Off? Contemporary Artificial intelligence Performance Evaluation on Adult Medicine Long Cases
Gao, C.; Bellinge, J.; El-Masri, S.; Chim, I.; Sorich, M.; Seth, I.; Gorcilov, J.; Lim, M.; McCoy, L.; Vanlint, A.; Lim, L.; Stranks, J.; Zannettino, A.; Maddison, J.; Bacchi, S.
Show abstract
The Royal Australasian College of Physicians (RACP) "Long Cases" assess a trainees clinical reasoning beyond what is tested in multiple-choice questions, which large language models (LLMs) have already demonstrated proficiency in. This study evaluated a LLMs ability to perform a "Long Case" assessment, including history-taking, case presentation, and answering examiner questions. The LLM achieved passing consensus scores of 4-5 out of 6 on five cases, suggesting potential for LLMs in complex clinical evaluations.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Large language models for generating medical examinations: systematic review 94%
- Evaluation of Statistical Illiteracy in Latin American Clinicians and of the Efficacy of a 10-Hour Course 94%
- Medical students' perceptions towards artificial intelligence in education and practice: A multinational, multicenter cross-sectional study 93%
Similar papers in this journal
- Performance of ChatGPT, GPT-4, and Google Bard on a Neurosurgery Oral Boards Preparation Question Bank 94%
- Performance of ChatGPT and GPT-4 on Neurosurgery Written Board Examinations 94%
- A Porcine Model of Peripheral Nerve Injury Enabling Ultra-Long Regenerative Distances: Surgical Approach, Recovery Kinetics, and Clinical Relevance 83%
Similar papers in this journal
- Bridging the Literacy Gap for Surgical Consents: An AI-Human Expert Collaborative Approach 93%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 92%
- Utilization of Generative AI-drafted Responses for Managing Patient-Provider Communication 91%
Similar papers in this journal
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 93%
- Prohibiting Babel - A call for professional remote interpreting services in pre-operation anaesthesia information 92%
- Evaluation of Large Language Models in Medical Examinations:A Scoping Review Protocol 92%
Similar papers in this journal
- AI-Simulated Clinical Consultations: Assessing the Potential of ChatGPT to Support Medical Training 94%
- General population screening for type 1 diabetes using islet autoantibodies at the preschool vaccination visit: a proof-of-concept study (the T1Early study) 87%
- ‘Admissions to paediatric medical wards with a primary mental health diagnosis: a systematic review of the literature’ 87%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.