Retrieval-Augmented Generation for Extracting CHA2DS2VASc Features from Unstructured Clinical Notes in Patients with Atrial Fibrillation
Adejumo, P.; Thangaraj, P. M.; Vasisht Shankar, S.; Dhingra, L. S.; Aminorroaya, A.; Khera, R.
Show abstract
ImportanceStandardized assessment of clinical quality measures from electronic health records (EHRs) is challenging because information is fragmented across structured and unstructured data, and due to low interoperability across systems. Traditionally, extracting this information requires manual EHR abstraction, a time-consuming and expensive process that also limits real-time care quality improvement. Objective: To evaluate whether a data format-agnostic retrieval-augmented generation-enabled large language model (RAG-LLM) can accurately abstract clinical variables from heterogeneous structured and unstructured EHR data. Design, Setting, and ParticipantsRetrospective cross-sectional study assessing stroke and bleeding risk in patients with atrial fibrillation (AF) from two health systems. We developed a RAG-LLM model to extract CHA DS-VASc and HAS-BLED risk factors from tabular data and clinical documentation. The framework was validated on 300 expert-annotated patient records (200 from Yale New Haven Health System [YNHHS] and 100 from the Medical Information Mart for Intensive Care [MIMIC-IV]). The system was deployed on two large cohorts: 104,204 patients with AF from YNHHS (2013-2024) and 13,117 from MIMIC-IV (2008-2022). We compared anticoagulation recommendations derived from RAG-LLM with those based on traditional structured data abstraction. ExposuresUse of a RAG-LLM model to abstract stroke and bleeding risk factors from structured and unstructured EHR data. Main Outcomes and MeasuresAccuracy of RAG-LLM-based risk factor abstraction against expert annotation. Secondary outcomes included efficiency, cross-cohort generalizability, and impact on anticoagulation eligibility based on risk stratification. ResultsIn the validation cohort (mean age 74.8 years, 42.7% female), RAG-LLM demonstrated superior performance across all metrics compared with structural data abstraction. For individual CHA DS-VASc components, accuracy ranged from 0.94-1.00 (YNHHS) and 0.89-1.00 (MIMIC-IV) versus 0.66-0.92 (YNHHS) and 0.44-0.97 (MIMIC-IV) for structured data, which was similar for HAS-BLED (0.94-1.00 and 0.89-1.00 vs 0.66-0.94 and 0.44-0.97). In the deployment study, among 3,207 patients classified as low/intermediate stroke risk with structured data, 62.1% (1,993) were reclassified as high risk with RAG-LLM and would become eligible for anticoagulation. Similarly, 5.5% of those classified as low bleeding risk by structured data were reclassified as high risk, substantially refining contraindication assessment. ConclusionsA multimodal RAG-LLM accurately abstracts clinical variables from structured and unstructured EHR data to improve stroke and bleeding risk assessments in patients with AF, enhancing identification of appropriate anticoagulation candidates.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Comparative Analysis of Privacy-Preserving Large Language Models For Automated Echocardiography Report Analysis 95%
- Biometric Contrastive Learning for Data-Efficient Deep Learning from Electrocardiographic Images 94%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 93%
Similar papers in this journal
- EchoGraph System for Automated Quality Assessment of Echocardiography Reports 94%
- Cohort Design and Natural Language Processing to Reduce Bias in Electronic Health Records Research: The Community Care Cohort Project 94%
- Aggregating Multiple Real-World Data Sources using a Patient-Centered Health Data Sharing Platform: an 8-week Cohort Study among Patients Undergoing Bariatric Surgery or Catheter Ablation of Atrial Fibrillation 94%
Similar papers in this journal
- Natural Language Processing for the Ascertainment and Phenotyping of Left Ventricular Hypertrophy and Hypertrophic Cardiomyopathy on Echocardiogram Reports 95%
- Risk of Cardiovascular Events after Covid-19: a double-cohort study 92%
- Ticagrelor vs Clopidogrel: the Impact of Platelet Inhibition on Cerebrovascular Microembolic Events during TAVR 92%
Similar papers in this journal
- Automated Diagnostic Reports from Images of Electrocardiograms at the Point-of-Care 95%
- International Evaluation Of An Artificial Intelligence-Powered Ecg Model Detecting Occlusion Myocardial Infarction 94%
- Simple Models Versus Deep Learning in Detecting Low Ejection Fraction From The Electrocardiogram 94%
Similar papers in this journal
- An assessment of the value of deep neural networks in genetic risk prediction for surgically relevant outcomes 95%
- Leveraging Machine Learning for Enhanced and Interpretable Risk Prediction of Venous Thromboembolism in Acute Ischemic Stroke Care 94%
- ChatGPT Provides Inconsistent Risk-Stratification of Patients With Atraumatic Chest Pain 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.