Real-World Performance of Large Language Models in Emergency Department Chest Pain Triage
meng, x.; Tang, Y.-D.
Show abstract
BackgroundLarge Language Models (LLMs) are increasingly being explored for medical applications, particularly in emergency triage where rapid and accurate decision-making is crucial. This study evaluates the diagnostic performance of two prominent Chinese LLMs, "Tongyi Qianwen" and "Lingyi Zhihui," alongside a newly developed model, MediGuide-14B, comparing their effectiveness with human medical experts in emergency chest pain triage. MethodsConducted at Peking University Third Hospitals emergency centers from June 2021 to May 2023, this retrospective study involved 11,428 patients with chest pain symptoms. Data were extracted from electronic medical records, excluding diagnostic test results, and used to assess the models and human experts in a double-blind setup. The models performances were evaluated based on their accuracy, sensitivity, and specificity in diagnosing Acute Coronary Syndrome (ACS). Findings"Lingyi Zhihui" demonstrated a diagnostic accuracy of 76.40%, sensitivity of 90.99%, and specificity of 70.15%. "Tongyi Qianwen" showed an accuracy of 61.11%, sensitivity of 91.67%, and specificity of 47.95%. MediGuide-14B outperformed these models with an accuracy of 84.52%, showcasing high sensitivity and commendable specificity. Human experts achieved higher accuracy (86.37%) and specificity (89.26%) but lower sensitivity compared to the LLMs. The study also highlighted the potential of LLMs to provide rapid triage decisions, significantly faster than human experts, though with varying degrees of reliability and completeness in their recommendations. InterpretationThe study confirms the potential of LLMs in enhancing emergency medical diagnostics, particularly in settings with limited resources. MediGuide-14B, with its tailored training for medical applications, demonstrates considerable promise for clinical integration. However, the variability in performance underscores the need for further fine-tuning and contextual adaptation to improve reliability and efficacy in medical applications. Future research should focus on optimizing LLMs for specific medical tasks and integrating them with conventional medical systems to leverage their full potential in real-world settings.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 96%
- Development and Prospective Implementation of a Large Language Model based System for Early Sepsis Prediction 95%
- CT-based Rapid Triage of COVID-19 Patients: Risk Prediction and Progression Estimation of ICU Admission, Mechanical Ventilation, and Death of Hospitalized Patients 95%
Similar papers in this journal
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 96%
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 95%
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 94%
Similar papers in this journal
- SymScore: Machine Learning Accuracy Meets Transparency in a Symbolic Regression-Based Clinical Score Generator 95%
- AI-MET: A Deep Learning-based Clinical Decision Support System for Distinguishing Multisystem Inflammatory Syndrome in Children from Endemic Typhus 94%
- Refining LLMs Outputs with Iterative Consensus Ensemble (ICE) 94%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.