Back

Not Such a Long Way Off? Contemporary Artificial intelligence Performance Evaluation on Adult Medicine Long Cases

Gao, C.; Bellinge, J.; El-Masri, S.; Chim, I.; Sorich, M.; Seth, I.; Gorcilov, J.; Lim, M.; McCoy, L.; Vanlint, A.; Lim, L.; Stranks, J.; Zannettino, A.; Maddison, J.; Bacchi, S.

2025-01-27 medical education
10.1101/2025.01.25.25321118 medRxiv
Show abstract

The Royal Australasian College of Physicians (RACP) "Long Cases" assess a trainees clinical reasoning beyond what is tested in multiple-choice questions, which large language models (LLMs) have already demonstrated proficiency in. This study evaluated a LLMs ability to perform a "Long Case" assessment, including history-taking, case presentation, and answering examiner questions. The LLM achieved passing consensus scores of 4-5 out of 6 on five cases, suggesting potential for LLMs in complex clinical evaluations.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.