Consistency in Large Language Models Ensures Reliable Patient Feedback Classification
Loi, Z.; Morquin, D.; Derzko, F. X.; Corbier, X.; Gauthier, S.; Moniez, L.; Prin Lombardo, E.; Mercier, G.; Yauy, K.
Show abstract
BackgroundPatient satisfaction feedback is crucial for hospital service quality, but manual reviews are not possible due to their time-consumption, and traditional natural language processing methods remain inadequate. Large Language Models (LLMs) show promise but are prone to logical hallucinations--fabricated or illogical outputs that limit their reliability (inconsistent performance across repeated uses) and validity (explainability into clinical contexts) in healthcare. ObjectiveThis study aimed to evaluate the Self-Logical Consistency Assessment (SLCA), an original method designed to enhance LLM feedback classification reliability by enforcing a logically-structured chain of thought. MethodsSLCA uses two validation steps: self-consistency (identifying the most coherent response) and logical consistency (ensuring alignment with the original statement and expert classifications). We evaluated SLCA using GPT-4 and Llama-3.1 405B on 12,600 classifications from 100 patient feedback samples to assess logical hallucinations, and tested its performance on a 49,140-classification benchmark derived from 1,170 feedbacks. ResultsSLCA reduced logical hallucinations among detected categories from 15.80% (168/1063) to 0.51% (4/786) with GPT-4 and from 7.17% (51/711) to 1.67% (10/599) with Llama-3.1, with residual errors confined to the emergency feedback category. On the benchmark, SLCA achieved precision-recall scores of 0.86-0.78 for GPT-4 and 0.84-0.58 for Llama-3.1. These results demonstrate SLCAs ability to achieve human-level performance across LLMs. ConclusionsSLCA offers a zero-shot, scalable, explainable solution for improving LLM classification reliability in healthcare. Its capacity to enhance performance without fine-tuning positions it as a valuable tool for analyzing patient feedback and supporting hospital service quality improvement. What is already known on this topicFree-text patient feedback labeling is crucial for healthcare system improvement. Classification by hand is impractical due to its time-consumption. Large language models (LLMs) can out-perform traditional NLP for this task, but their clinical use is limited by inconsistent predictions and "logical hallucinations" that undermine explainability and trust. What this study addsThe Self-Logical Consistency Assessment (SLCA) framework, which couples self-consistency with a novel logical-consistency check, almost eradicates hallucinations (from 15.8 % to 0.5 % with GPT-4) while reaching human-level precision (86%) and better exhaustivity (recall +14%). How this study might affect research, practice or policySLCA offers a scalable, explainable and data-sovereign pathway for hospitals and regulators to adopt LLMs in routine patient-experience monitoring, and it provides a transferable template for evaluating AI safety in other clinical-text tasks. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=142 HEIGHT=200 SRC="FIGDIR/small/24310210v4_ufig1.gif" ALT="Figure 1"> View larger version (31K): org.highwire.dtl.DTLVardef@147cab0org.highwire.dtl.DTLVardef@4c19aeorg.highwire.dtl.DTLVardef@2a20ecorg.highwire.dtl.DTLVardef@1d7a082_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOGraphical abstractC_FLOATNO C_FIG
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 97%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 95%
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 95%
Similar papers in this journal
Similar papers in this journal
- Optimising supervised machine learning algorithms predicting cigarette cravings and lapses for a smoking cessation just-in-time adaptive intervention (JITAI) 94%
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 93%
- A method for rapid machine learning development for data mining with Doctor-In-The-Loop 93%
Similar papers in this journal
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 94%
- Addressing Label Noise for Electronic Health Records: Insights from Computer Vision for Tabular Data 93%
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.