Building and validating 5-feature models to predict preeclampsia onset time from electronic health record data
Ballard, H. K.; Yang, X.; Mahadevan, A. D.; Garmire, L. K.; Lemas, D. J.
Show abstract
BackgroundPreeclampsia is a potentially fatal complication during pregnancy, characterized by high blood pressure and presence of proteins in the urine. Due to its complexity, prediction of preeclampsia onset is often difficult and inaccurate. MethodsThis study aims to create quantitative models to predict the onset gestational age of preeclampsia using electronic health records. We retrospectively collected 1178 preeclamptic pregnancy records from the University of Michigan Health System(UM) as the discovery cohort, and 881 records from the University of Florida Health System(UF) as the validation cohort. We constructed two Cox-proportional hazards models with Lasso regularization: one baseline model utilizing maternal and pregnancy characteristics, and the other full model with additional lab results, vital signs, and medications in the first 20 weeks of pregnancy. We built the models using 80% of the UM data and subsequently tested them on the remaining 20% UM data and validated with UF data. We further stratified the patients into high and low risk groups for preeclampsia onset risk assessment. FindingsThe baseline model reached C-indices of 0{middle dot}64 and 0{middle dot}61 in the 20% UM testing data and the UF validation data, respectively, while the full model increased these C-indices to 0{middle dot}69 and 0{middle dot}61 respectively. Both the baseline and full models contain five selective features, among which number of fetuses in the pregnancy, hypertension and parity are shared between the two models with similar hazard ratios. In the baseline model, history of complicated type II diabetes and a mood/anxiety disorder during the first 20 weeks of pregnancy were important. In the full model, maximum diastolic blood pressure in early pregnancy was the predominant feature. InterpretationElectronic health record data provide useful information to predict gestational age of preeclampsia onset. Stratification of the cohorts using five-predictor Cox-PH models provide clinicians with convenient tools to assess the patients onset time of preeclampsia. FundingThis study was supported by grants through the NIEHS, NICHD, NIDDK, and NCATS.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Preeclampsia prediction with maternal and paternal polygenic risk scores: the TMM BirThree Cohort Study 96%
- Widely accessible prognostication using medical history for fetal growth restriction and small for gestational age in nationwide insured women 95%
- Genetic Predictors of Blood Pressure Traits are Associated with Preeclampsia 94%
Similar papers in this journal
- Improving Pre-eclampsia Risk Prediction by Modeling Individualized Pregnancy Trajectories Derived from Routinely Collected Electronic Medical Record Data 98%
- Zero-shot Interpretable Phenotyping of Postpartum Hemorrhage Using Large Language Models 94%
- Cohort Design and Natural Language Processing to Reduce Bias in Electronic Health Records Research: The Community Care Cohort Project 91%
Similar papers in this journal
- Automated Interpretable Discovery of Heterogeneous Treatment Effectiveness: A Covid-19 Case Study 91%
- Disease Network Delineates the Disease Progression Profile of Cardiovascular Diseases 90%
- Natural language processing for scalable feature engineering and ultra-high-dimensional confounding adjustment in healthcare database studies 89%
Similar papers in this journal
Similar papers in this journal
- Prediction of Preeclampsia from Clinical and Genetic Risk Factors in Early and Late Pregnancy Using Machine Learning and Polygenic Risk Scores 96%
- Fetal sexual dimorphism and preeclampsia among twin pregnancies 93%
- Evidence from human placenta, ER-stressed trophoblasts and transgenic mice links transthyretin proteinopathy to preeclampsia 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.