Assessing the Effectiveness of ChatGPT as a Clinical Trainee: A Study on the Diagnostic Value of Large Language Models in a Complex Clinical Environment
Craig, D.; Nugent, C.
Show abstract
We tested the performance of Chat Generative Pre-trained Transformer (ChatGPT) in the role of a trainee clinician (Specialist Registrar or Resident) undergoing direct assessment by a human supervising specialist clinician (Consultant or Attending). The session consisted of a hospital ward round scenario presented to three versions of ChatGPT, namely OpenAI ChatGPT-3.5, Bing ChatGPT-4 and OpenAI ChatGPT-4. A specific test of memory and context was included via an end-of-teaching educator feedback exercise. Only OpenAI ChatGPT-4 provided responses comparable to the standard a trainee might offer during progress towards completion of training and specialist accreditation. Bing ChatGPT-4 responded with several clinically dubious statements, often in a repetitive and detached way, and was unable to retain awareness of the purpose of the session and the identities of participants.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Harnessing the Open Access Version of ChatGPT for Enhanced Clinical Opinions 95%
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 94%
- Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study 94%
Similar papers in this journal
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 95%
- The performance of national COVID-19 ‘Symptom Checkers’: A comparative case simulation study 94%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 93%
Similar papers in this journal
- A method for rapid machine learning development for data mining with Doctor-In-The-Loop 96%
- Evaluating user experience with immersive technology in simulation-based education: a modified Delphi study with qualitative analysis 94%
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 94%
Similar papers in this journal
- Improving emergency department patient-doctor conversation through an artificial intelligence symptom taking tool: an action-oriented design pilot study 94%
- The potential for digital patient symptom recording through symptom assessment applications to optimize patient flow and reduce waiting times in Urgent Care Centers: a simulation study 93%
- Exploring Patient and Staff Experiences of Video Consultations During COVID-19 in an English Outpatient Care Setting: Secondary Data Analysis of Routinely Collected Feedback Data 93%
Similar papers in this journal
- Electronic prescribing systems as tools to improve patient care: a learning health systems approach to increase guideline concordant prescribing for venous thromboembolism prevention 93%
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 93%
- Optimized Feature Selection and Advanced Machine Learning for Stroke Risk Prediction in Revascularized Coronary Artery Disease Patients 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.