The Scope and Limitations of Extant Research into ChatGPT as a Tool for Patient Education: Systematic Review
Dale, R.; Cheng, M.; Casselman Pines, K.; Currie, M.
Show abstract
BackgroundChat Generative Pre-Trained Transformer (ChatGPT), a large language model (LLM) developed by OpenAI, has been extensively studied and embraced by medical researchers since its public release in November 2022. In addition to its benefits in generating summaries and predictive diagnostics, it has been proposed as a patient education tool. Existing research cites ChatGPTs potential to increase information availability and accessibility, improve the efficiency of clinical practice, and its high-quality responses to clinical questions as reasons to consider its adoption. ObjectiveWe assessed literature from PubMed on the quality and consistency of ChatGPT responses across medical specialties, evaluated the comprehensiveness of current research, and identified areas for future research. MethodsWe searched PubMed for published articles evaluating the consistency, reliability, and ethics of ChatGPT in patient education. Following the PRISMA guideline, we conducted a systematic literature review of the 567 retrieved records. After title, abstract, and full-text screens, 123 relevant records were included and synthesized in this review. ResultsWe found a lack of consensus among ChatGPT studies. The model accuracy suffers from infrequent updates and generation of misleading information (hallucinations), and it lacks knowledge of current clinical guidelines. The consistency of the model falls short due to its sensitivity to prompt design and fine-tuning through user interaction, making ChatGPT research results almost impossible to peer review and validate. Relying on ChatGPT for clinical information risks spreading misinformation, disrupting trust in the medical system, and disobeying the principles of patient-centered care. This also shifts the focus of patient education from shared decision-making and information-building to information-giving, deviating from the objectives of Health Communication and Health Literacy outlined in Healthy People 2030. ConclusionsWe caution against the acceptance of ChatGPT as a patient education tool and encourage future efforts to incorporate community perspectives in AI research.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Harnessing the Open Access Version of ChatGPT for Enhanced Clinical Opinions 94%
- Development and preliminary testing of Health Equity Across the AI Lifecycle (HEAAL): A framework for healthcare delivery organizations to mitigate the risk of AI solutions worsening health inequities 93%
- Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study 93%
Similar papers in this journal
- Simulated Misuse of Large Language Models and Clinical Credit Systems 95%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 94%
- Can co-designed educational interventions help consumers think critically about asking ChatGPT health questions? Results from a randomised-controlled trial 94%
Similar papers in this journal
- Cracking the Code: A Scoping Review to Unite Disciplines in Tackling Legal Issues in Health Artificial Intelligence 93%
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 92%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 92%
Similar papers in this journal
- A scoping review of fair machine learning techniques when using real-world data 93%
- Development of a Post-Acute Sequelae of COVID-19 (PASC) Symptom Lexicon Using Electronic Health Record Clinical Notes 92%
- Demonstrating the Consequences of Learning Missingness Patterns in Early Warning Systems for Preventative Health Care: A Novel Simulation and Solution 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.