Back

Machine Learning-Supported Efficient VTE Risk Assessment using Routinely Collected Electronic Health Record Data

Li, Z.; Yagis, E.; Riad, A.; Windrath-Carr, O.; Arribas, M.; Sodiq, T.; Goldsmith, K.; Glampson, B.; Flott, K.; Haji, G.; Khan, Z.; Baker, C.; Mayer, E. K.

2026-08-19 health informatics
10.64898/2026.08.18.26360687 medRxiv
Show abstract

Venous thromboembolism (VTE) is a leading cause of preventable inpatient mortality, while the real-world performance of mandated risk assessment and the potential for automating using electronic health record (EHR) data remain unclear. We analysed 577,904 admissions and 726,896 VTE assessment forms across five NHS hospitals between 2015 and 2025 to evaluate assessment completion, concordance with structured EHR data, clinical validity, and feasibility of EHR-based automation assisted by machine learning. Overall completion was high (96.7%), and timely completion improved from 47.4% in 2015 to 90.5% in 2024. Agreement between forms and EHR data was good for common risk factors, but low-prevalence variables were often under-documented in the forms. Despite these discrepancies, form-derived thrombosis risk was associated with increased VTE incidence (OR 3.31, 95% CI 2.81-3.90). Machine learning models using first-14-hour EHR data achieved discrimination comparable to clinician-recorded variables (AUROC 0.709 vs 0.704), supporting real-time EHR-integrated assessment pre-population and decision support.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.