The Pulse of Artificial Intelligence in Cardiology: A Comprehensive Evaluation of State-of-the-art Large Language Models for Potential Use in Clinical Cardiology
Novak, A.; Rode, F.; Lisicic, A.; Nola, I. A.; Zeljkovic, I.; Pavlovic, N.; Manola, S.
Show abstract
IntroductionOver the past two years, the use of Large Language Models (LLMs) in clinical medicine has expanded significantly, particularly in cardiology, where they are applied to ECG interpretation, data analysis, and risk prediction. This study evaluates the performance of five advanced LLMs--Google Bard, GPT-3.5 Turbo, GPT-4.0, GPT-4o, and GPT-o1-mini--in responding to cardiology-specific questions of varying complexity. MethodsA comparative analysis was conducted using four test sets of increasing difficulty, encompassing a range of cardiovascular topics, from prevention strategies to acute management and diverse pathologies. The models responses were assessed for accuracy, understanding of medical terminology, clinical relevance, and adherence to guidelines by a panel of experienced cardiologists. ResultsAll models demonstrated a foundational understanding of medical terminology but varied in clinical application and accuracy. GPT-4.0 exhibited superior performance, with accuracy rates of 92% (Set A), 88% (Set B), 80% (Set C), and 84% (Set D). GPT-4o and GPT-o1-mini closely followed, surpassing GPT-3.5 Turbo, which scored 83%, 64%, 67%, and 57%, and Google Bard, which achieved 79%, 60%, 50%, and 55%, respectively. Statistical analyses confirmed significant differences in performance across the models, particularly in the more complex test sets. While all models demonstrated potential for clinical application, their inability to reference ongoing clinical trials and some inconsistencies in guideline adherence highlight areas for improvement. ConclusionLLMs demonstrate considerable potential in interpreting and applying clinical guidelines to vignette-based cardiology queries, with GPT-4.0 leading in accuracy and guideline alignment. These tools offer promising avenues for augmenting clinical decision-making but should be used as complementary aids under professional supervision.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Validating a Clinical Decision Support System for Palliative Care using healthcare professionals’ insights 95%
- A digital self-care intervention for Ugandan patients with heart failure and their clinicians: User-centred design and usability study 94%
- How suitable are clinical vignettes for the evaluation of symptom checker apps? A test theoretical perspective 93%
Similar papers in this journal
- The potential for digital patient symptom recording through symptom assessment applications to optimize patient flow and reduce waiting times in Urgent Care Centers: a simulation study 94%
- Remote patient monitoring and digital therapeutics in heart failure: lessons from the Continuum pilot study 93%
- Improving emergency department patient-doctor conversation through an artificial intelligence symptom taking tool: an action-oriented design pilot study 93%
Similar papers in this journal
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 93%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 92%
- Utilization of Generative AI-drafted Responses for Managing Patient-Provider Communication 92%
Similar papers in this journal
- Identification of Myocardial Infarction (MI) Probability from Imbalanced Medical Survey Data: An Artificial Neural Network (ANN) with Explainable AI (XAI) Insights 92%
- Improving irregular temporal modeling by integrating synthetic data to the electronic medical record using conditional GANs: a case study of fluid overload prediction in the intensive care unit 92%
- Refining LLMs Outputs with Iterative Consensus Ensemble (ICE) 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.