Development of an LLM Pipeline Surpassing Physicians in Cardiovascular Risk Score Calculation
Roeschl, T.; Hoffmann, M.; Unbehaun, A.; Dreger, H.; Hindricks, G.; Falk, V.; Balicer, R.; Tanacli, R.; Hohendanner, F.; Meyer, A.
Show abstract
BackgroundRisk scores are essential to evidence-based cardiovascular care, but manual calculation is labor-intensive and error-prone. Large language models (LLMs) could automate this process, yet LLMs are limited by their propensity for calculation errors and factual hallucinations. Pipelines that separate LLM-based data extraction from deterministic score computation may improve reliability and transparency. MethodsWe conducted a retrospective diagnostic study at a quaternary heart center in Germany (January 2020 - July 2023). Patients with atrial fibrillation (n=179) from an ablation registry and patients with severe aortic stenosis (n=76) evaluated by a heart team were included. Five LLMs (DeepSeek-R1, Qwen3, GPT-4 Turbo, Llama 3.1, and PaLM 2) were tested in standalone and pipeline configurations to compute HAS-BLED, CHA2DS2-VASc, and EuroSCORE II scores from routine clinical reports. Accuracy was assessed by comparing predictions to expert-adjudicated ground truth, using root mean squared error (RMSE), Krippendorffs alpha for categorical agreement, and calibration analysis. ResultsPipeline-generated scores showed substantially higher agreement with expert adjudication than standalone LLMs and treating clinicians (mean Krippendorffs alpha: 0.79 vs 0.32 vs 0.31) and demonstrated superior calibration. The Qwen3-based pipeline, achieved the highest accuracy with lower RMSEs than clinicians for HAS-BLED (0.20 vs 0.87), CHA2DS2-VASc (0.53 vs 1.08), and EuroSCORE II (1.99 vs 2.05). ConclusionLLM-based pipelines enable accurate, well-calibrated, and scalable cardiovascular risk score computation from unstructured real-world clinical data, outperforming clinicians and standalone LLMs with the potential to reduce clinician workload and support evidence-based care.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Biometric Contrastive Learning for Data-Efficient Deep Learning from Electrocardiographic Images 96%
- A Comparative Analysis of Privacy-Preserving Large Language Models For Automated Echocardiography Report Analysis 96%
- Learning Decision Thresholds for Risk-Stratification Models from Aggregate Clinician Behavior 92%
Similar papers in this journal
- Detecting QT prolongation From a Single-lead ECG With Deep Learning 92%
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 92%
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 92%
Similar papers in this journal
- Explaining Deep Neural Networks for Knowledge Discovery in Electrocardiogram Analysis 93%
- Evaluation of Domain Generalization and Adaptation on Improving Model Robustness to Temporal Dataset Shift in Clinical Medicine 93%
- Opportunistic Assessment of Ischemic Heart Disease Risk Using Abdominopelvic Computed Tomography and Medical Record Data: a Multimodal Explainable Artificial Intelligence Approach 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.