AI-Simulated Clinical Consultations: Assessing the Potential of ChatGPT to Support Medical Training
Saggar, A.; Dimitrova, V.; Sarikaya, D.; Hogg, D. C.; Darling, J. C.
Show abstract
BackgroundSimulated medical scenarios are useful for evaluating and developing clinical competencies but scheduling them is expensive and time-consuming. Large language models (LLMs) show promise in role-playing tasks. We investigated the fidelity with which ChatGPT can mimic patients, clinicians and examiners in educational settings. ObjectiveTo determine the realism with which ChatGPT can portray patient, doctor and examiner roles, and the utility of these agents in clinical education. MethodWe selected four paediatric scenarios from mock OSCEs and set up separate patient, doctor and examiner ChatGPT agents for each. The patient and doctor agents conversed with each other in written format. The examiner agent marked the doctor agent based on this conversation. Patients and clinicians familiar with the OSCE assessed the dialogues. ResultsThe patient agent was judged to be true to character most of the time and good at expressing emotion. The doctor agent was reported to be an effective communicator but occasionally used jargon. Both agents tended to produce repetitive responses which undermined realism. The examiner agent had good correlation with human clinicians. There was moderate support for using the simulated interactions for educational purposes. ConclusionAlthough the realism of the agents can be improved, ChatGPT can generate plausible proxies of participants in medical scenarios and could be useful for complementing standardised patient (SP)-based training. KEY MESSAGESO_ST_ABSWhat is already known on this topicC_ST_ABSO_LILLM-based agents show promise for portraying clinical roles and supporting simulation-based learning. Doctor agents provide correct diagnoses most of the time, while patient agents can accurately relay role information such as medical history or symptoms. C_LI What this study addsO_LIThere is scope for improvement in the realism and authenticity of the conversations produced by GPT patient and doctor agents. Notable issues included a tendency to produce repetitive and verbose responses, and an inability to accurately convey the hesitation shown by real patients. C_LIO_LIDisparities observed between (human) patient and clinician assessment for the GPT agents suggest that diverse viewpoints are needed to fully capture the experiential learning associated with clinical communication. C_LIO_LIHow this study might affect research, practice or policy C_LIO_LILow fidelity of GPT simulations for difficult or challenging medical scenarios necessitates human oversight and correction for AI deployed in educational settings. C_LIO_LIThe impact of AI on medical education is likely to increase in the future, which necessitates promoting AI literacy among educators and students. C_LI
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Large language models for generating medical examinations: systematic review 95%
- Improving capacity for advanced training in obstetric surgery: Evaluation of a blended learning approach 94%
- Medical students' perceptions towards artificial intelligence in education and practice: A multinational, multicenter cross-sectional study 93%
Similar papers in this journal
- Evaluating user experience with immersive technology in simulation-based education: a modified Delphi study with qualitative analysis 94%
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 94%
- Key topics in pandemic health risk communication: A qualitative study of expert opinions and knowledge 93%
Similar papers in this journal
- Exploring Patient and Staff Experiences of Video Consultations During COVID-19 in an English Outpatient Care Setting: Secondary Data Analysis of Routinely Collected Feedback Data 93%
- Design and Formative Evaluation of a Voice-based Virtual Coach for Problem-Solving Treatment 93%
- Improving emergency department patient-doctor conversation through an artificial intelligence symptom taking tool: an action-oriented design pilot study 92%
Similar papers in this journal
Similar papers in this journal
- Assessing ChatGPT’s Mastery of Bloom’s Taxonomy using psychosomatic medicine exam questions 94%
- Tracking private WhatsApp discourse about COVID-19: A longitudinal infodemiology study in Singapore 92%
- Artificial Intelligence (AI)-based Chatbots in Promoting Health Behavioral Changes: A Systematic Review 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.