Leveraging Large Language Models to Develop an Interpretable Prediction Model for Postpartum Hemorrhage Prior to the Onset of Labor
Woo, E. G.; Zighelboim, I.; Gifford, T.; Bell, J. G.; Milthorpe, H.; Alsentzer, E.; Longman, R. E.; Tolosa, J. E.; Beaulieu-Jones, B. K.
Show abstract
ObjectiveTo evaluate whether large language models (LLMs) applied to prenatal clinical notes can predict postpartum hemorrhage (PPH) prior to the onset of labor and to compare model performance across outcome definitions, including a novel intervention-based definition. MethodsWe conducted a retrospective cohort study of 19,992 deliveries within a large regional health network. Two outcome definitions for PPH were used: estimated or quantitative blood loss (EBL/QBL) extracted from clinical notes, and a clinical intervention-based definition (cPPH) incorporating transfusion, uterotonics, Bakri balloon, or hysterectomy. We evaluated three approaches for PPH prediction: (1) supervised machine learning using structured electronic medical record data; (2) direct prediction using a fine-tuned LLM applied to clinical notes; and (3) interpretable models using LLM-extracted features combined with structured data. Model performance was evaluated using area under the receiver operating characteristic curve (AUROC) on a temporally held-out test set. ResultsThe LLM-based direct prediction model achieved the highest performance for both PPH definitions (AUROC 0.79-0.80), followed by interpretable models combining LLM-extracted features with structured data (AUROC 0.76-0.78). Models using only structured data performed worse (AUROC 0.65-0.71). The LLM-extracted features approach identified 47 significant predictors, including established risk factors such as multiple gestation and previous cesarean delivery. Demographic differences were observed between PPH definitions: mothers who met only the cPPH definition had lower gestational age and higher rates of cesarean delivery compared to those meeting only the EBL/QBL definition. ConclusionThese findings highlight the potential of LLM-based approaches for enhancing PPH risk stratification, with the feature extraction method offering a promising balance between predictive performance and clinical utility. Integrating these methods into clinical workflows could improve early detection and guide targeted preventive interventions.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Zero-shot Interpretable Phenotyping of Postpartum Hemorrhage Using Large Language Models 98%
- Improving Pre-eclampsia Risk Prediction by Modeling Individualized Pregnancy Trajectories Derived from Routinely Collected Electronic Medical Record Data 96%
- Continuous-Time and Dynamic Suicide Attempt Risk Prediction with Neural Ordinary Differential Equations 93%
Similar papers in this journal
- A comprehensive digital phenotype for postpartum hemorrhage 95%
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 92%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 92%
Similar papers in this journal
Similar papers in this journal
- Widely accessible prognostication using medical history for fetal growth restriction and small for gestational age in nationwide insured women 97%
- Large Language Models Improve the Identification of Emergency Department Visits for Symptomatic Kidney Stones 93%
- Scalable Incident Detection via Natural Language Processing and Probabilistic Language Models 93%
Similar papers in this journal
- A Deep Learning Method to Detect Opioid Prescription and Opioid Use Disorder from Electronic Health Records 92%
- Predicting nutrition and environmental factors associated with female reproductive disorders using a knowledge graph and random forests 91%
- Predicting Prognosis in COVID-19 Patients using Machine Learning and Readily Available Clinical Data 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.