Multisite Real-World Validation of an Electronic Health Record-Integrated Generative Artificial Intelligence Tool for Venous Thromboembolism Risk Stratification
Baughman, D. J.; Liu, S.; Jee, S.; Young, C.; Knight, A. M.; Davis, A.; Yegnasubramanian, S.; Najjar, P.; Whitbread, J. J.; Ahumada, L.; Chused, A.; Haut, E. R.; Lau, B. D.; Sridharan, A.; Streiff, M.; Aziz, K. B.
Show abstract
Background: Guiding risk-appropriate inpatient thromboprophylaxis requires venous thromboembolism (VTE) risk stratification; however, reliable risk determination remains inconsistent in routine care. Health systems increasingly pilot artificial intelligence (AI) tools, yet few studies demonstrate rigorous evaluation in the context of a learning health system (LHS). We evaluated the performance of a pilot electronic health record (EHR)-integrated generative AI (GenAI) system, inHealth General Reasoner (iHGR), for VTE risk stratification versus clinician order set classifications and physician-adjudicated chart review. Methods: This multisite retrospective validation study included adult inpatient admissions at Johns Hopkins Medicine between June 21, 2025, and Dec 18, 2025 (checklist-based order set from June 21, 2025 - November 19, 2025, and clinician judgement-based order set from November 29 - December 18, 2025). From 758 eligible admissions, we randomly sampled 500 balanced by site and order set periods. iHGR and clinician-selected order set classifications were compared with the reference standard (RS). Primary outcomes were iHGR sensitivity and specificity. Secondary analyses compared the order sets with the same RS to evaluate workflow comparators and error patterns. Results: iHGR achieved 81.8% sensitivity (95% CI 77.3-85.6) and 70.9% specificity (63.6-77.3). The checklist-based order set had 61.3% sensitivity (53.7-68.5) and 86.2% specificity (77.4-91.9). The clinician judgement-based order set had 78.1% sensitivity (71.3-83.7) and 65.4% specificity (54.3-75.0). False-negative iHGR classifications were associated with missed narrative risk factors. Conclusion: iHGR showed higher sensitivity for VTE risk than checklist-based order sets and clinician judgement without introducing systematic bias. In silico evaluation of pilot AI systems within LHSs can identify clinically important performance trade-offs and implementation targets before operational scale-up. Narrative clinical data abstraction remained a key limitation, supporting the use of GenAI to support rather than supplant clinician judgement.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Real-Time Electronic Health Record Mortality Prediction During the COVID-19 Pandemic: A Prospective Cohort Study 94%
- Clinical Utility of Automatable Prediction Models for Improving Palliative and End-Of-Life Care Outcomes: Towards Routine Decision Analysis Before Implementation 94%
- Validation of a Derived International Patient Severity Algorithm to Support COVID-19 Analytics from Electronic Health Record Data 94%
Similar papers in this journal
- Bridging the Literacy Gap for Surgical Consents: An AI-Human Expert Collaborative Approach 93%
- Zero-shot Interpretable Phenotyping of Postpartum Hemorrhage Using Large Language Models 92%
- Development and Prospective Implementation of a Large Language Model based System for Early Sepsis Prediction 92%
Similar papers in this journal
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 92%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 90%
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 90%
Similar papers in this journal
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 94%
- Score for Emergency Risk Prediction (SERP): An Interpretable Machine Learning AutoScore–Derived Triage Tool for Predicting Mortality after Emergency Admissions 93%
- Diagnostic Codes in AI prediction models and Label Leakage of Same-admission Clinical Outcomes 90%
Similar papers in this journal
- Clinical Decision Support in Cardiovascular Medicine: Effectiveness, Implementation Barriers, and Regulation 93%
- Early initiation of prophylactic anticoagulation for prevention of COVID-19 mortality: a nationwide cohort study of hospitalized patients in the United States 92%
- Prone positioning of patients with moderate hypoxia due to COVID-19: A multicenter pragmatic randomized trial [COVID-PRONE] 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.