Machine Learning for Dynamic and Short-term Prediction of Preeclampsia Using Routine Clinical and Laboratory Data
Li, H.; Li, Y.; Zang, C.; Pan, W.; Yang, H. S.; Grossman, T. B.; zhao, Z.; Wang, F.
Show abstract
Preeclampsia (PE) is a leading cause of maternal and perinatal morbidity and mortality, yet its unpredictable onset and rapid progression hinder timely management. Existing prediction tools often rely on specialized biomarkers, static assessments, or limited study cohorts, impeding clinical utility and generalizability. We conducted a retrospective, multi-site cohort study including 58,839 pregnancies delivered at three NewYork-Presbyterian hospitals. Using routine information captured within the electronic health record (EHR), including blood pressure with other maternal characteristics, and routine laboratory tests, we developed extreme gradient boosting (XGBoost) based models to predict PE onset within 1-, 2-, and 4-week horizons across different gestational ages. Performance was assessed using nested cross-validation at the training site and externally validated through direct transfer, fine-tuning, and retraining strategies. Prediction accuracy increased from 28 to 34 weeks of gestational age, peaked at 34 weeks (AUC 0.863 at training; 0.808-0.834 at validation), declined at 38 weeks, and rebounded near delivery (AUC up to 0.890). Blood pressure was the most consistent predictor, while laboratory features such as albumin, alkaline phosphatase, and hematologic indices added value earlier, and demographic and obstetric factors gaining importance later. Dynamic short-term prediction of PE in late gestation is feasible using routine data. This pragmatic, scalable approach provides opportunities for early intervention and is adaptable across diverse healthcare settings.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Improving Pre-eclampsia Risk Prediction by Modeling Individualized Pregnancy Trajectories Derived from Routinely Collected Electronic Medical Record Data 98%
- Zero-shot Interpretable Phenotyping of Postpartum Hemorrhage Using Large Language Models 95%
- Continuous-Time and Dynamic Suicide Attempt Risk Prediction with Neural Ordinary Differential Equations 93%
Similar papers in this journal
- Clinical trial emulation can identify new opportunities to enhance the regulation of drug safety in pregnancy 92%
- Risk Factors for Spontaneous Preterm Birth are Mediated through Changes in Cervical Length 92%
- Genome-wide DNA methylation, imprinting, and gene expression in human placentas derived from Assisted Reproductive Technology 92%
Similar papers in this journal
- Widely accessible prognostication using medical history for fetal growth restriction and small for gestational age in nationwide insured women 96%
- Preeclampsia prediction with maternal and paternal polygenic risk scores: the TMM BirThree Cohort Study 94%
- SARS-CoV-2 (COVID-19) infection in pregnant women: characterization of symptoms and syndromes predictive of disease and severity through real-time, remote participatory epidemiology 94%
Similar papers in this journal
- Deep Learning Prediction of Biomarkers from Echocardiogram Videos 91%
- Transformer-based deep learning model for the diagnosis of suspected lung cancer in primary care based on electronic health record data 90%
- Predicting the functional effects of voltage-gated potassium channel missense variants with multi-task learning 90%
Similar papers in this journal
- Feasibility of Continuous Distal Body Temperature for Passive, Early Pregnancy Detection 92%
- Generalizability Challenges of Mortality Risk Prediction Models: A Retrospective Analysis on a Multi-center Database 91%
- Identification of physiological adverse events using continuous vital signs monitoring during paediatric critical care transport: a novel data-driven approach 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.