Mitigating Hallucinations in Large Language Models: A Comparative Study of RAG-enhanced vs. Human-Generated Medical Templates
Li, A.; Shrestha, R.; Jegatheeswaran, T.; Chan, H. O.; Hong, C.; Joshi, R.
Show abstract
The integration of Large Language Models (LLMs) is increasingly recognized for its potential to enhance various aspects of healthcare, including patient care, medical research, and education. The well-known LLM from Open AI: ChatGPT, a user-friendly GPT-4 based chatbot, has become increasingly popular. However, current limitations to LLMs, such as hallucinations, outdated information, and ethical and legal complications may pose significant risks to patients and contribute to the spread of medical disinformation. This study focuses on the application of Retrieval-Augmented Generation (RAG) to mitigate common limitations of LLMs like ChatGPT and assess its effectiveness in summarizing and organizing medical information. Up-to-date clinical guidelines were utilized as the source of information to create detailed medical templates. These were evaluated against human-generated templates by a panel of physicians, using Likert scales for accuracy and usefulness, and programmatically using BERTScores for textual similarity. The LLM templates scored higher on average for both accuracy and usefulness when compared to human-generated templates. BERTScore analysis further showed high textual similarity between ChatGPT- and Human-generated templates. These results indicate that RAG-enhanced LLM prompting can effectively summarize and organize medical information, demonstrating high potential for use in clinical settings.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 94%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 94%
- Network Graph Representation of COVID-19 Scientific Publications to Aid Knowledge Discovery 94%
Similar papers in this journal
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 96%
- What Do Clinicians Edit in Ambient AI-Drafted Clinical Documentation? A Qualitative Content Analysis 95%
- Usability of a Machine-Learning Clinical Order Recommender System Interface for Clinical Decision Support and Physician Workflow 94%
Similar papers in this journal
- The potential for digital patient symptom recording through symptom assessment applications to optimize patient flow and reduce waiting times in Urgent Care Centers: a simulation study 94%
- A Web-based, Mobile Responsive Application to Screen Healthcare Workers for COVID Symptoms: Descriptive Study 93%
- Improving emergency department patient-doctor conversation through an artificial intelligence symptom taking tool: an action-oriented design pilot study 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.