Back

Assessing the Effectiveness of ChatGPT as a Clinical Trainee: A Study on the Diagnostic Value of Large Language Models in a Complex Clinical Environment

Craig, D.; Nugent, C.

2023-11-22 health informatics
10.1101/2023.11.21.23298849 medRxiv
Show abstract

We tested the performance of Chat Generative Pre-trained Transformer (ChatGPT) in the role of a trainee clinician (Specialist Registrar or Resident) undergoing direct assessment by a human supervising specialist clinician (Consultant or Attending). The session consisted of a hospital ward round scenario presented to three versions of ChatGPT, namely OpenAI ChatGPT-3.5, Bing ChatGPT-4 and OpenAI ChatGPT-4. A specific test of memory and context was included via an end-of-teaching educator feedback exercise. Only OpenAI ChatGPT-4 provided responses comparable to the standard a trainee might offer during progress towards completion of training and specialist accreditation. Bing ChatGPT-4 responded with several clinically dubious statements, often in a repetitive and detached way, and was unable to retain awareness of the purpose of the session and the identities of participants.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.