Automated and interoperable methods for generalizable development of clinical machine-learning models for predicting neuromorbidity in critically ill children
Li, R.; Horvat, C. M.; Nourelahi, M.; Perez Claudio, E.; Hammett, J.; Wainwright, M. S.; Clark, R. S. B.; Au, A.; Hochheiser, H.
Show abstract
ObjectivesTo streamline the development of clinical machine learning (ML) models for predicting acute neurological morbidity in critically ill children by extending our prior work to create a standardized, reproducible, and scalable workflow leveraging Fast Healthcare Interoperability Resources (FHIR), cloud infrastructure, and automated ML tools. MethodsWe developed workflow for extracting, cleaning, and modeling pediatric intensive care unit (PICU) data, using 168 biomarkers from 7,403 encounters at an academic Childrens hospital between 2020 and 2024. Data were processed and stored in a compliant, secure cloud environment. We evaluated four feature sets: a baseline set from prior work, a complete set, a filtered set, and a light gradient boosting machine (LightGBM)-selected set for prediction of acquired neurological morbidity. Automated ML was used to train, validate, and deploy models, with performance assessed using the area under the receiver operating characteristics curve (AUROC), area under the precision recall curve (AUPRC), F1 score, calibration metrics, and Shapley additive (SHAP) values. A FHIR-based version of the pipeline was also implemented and evaluated on a 2020 subset of the cohort. ResultsFiltered and LightGBM-based feature sets achieved the highest predictive performance, with AUROCs of 0.90 (95% CI: [0.88-0.92]) for both, and AUPRCs of 0.68 (95% CI: [0.63-0.73]) and 0.67(95% CI: [0.62-0.72]), respectively. SHAP value analysis revealed consistent top features across models, with vital signs and key laboratory values prominently ranked. Models trained using FHIR-formatted data from a 2020 cohort (n = 1,339) demonstrated comparable performance to those built on the complete dataset, with an AUROC of 0.87 (95% CI: [0.81-0.93]). ConclusionsThis study demonstrates the feasibility of a cloud-compatible, standards-based approach to clinical ML model development. By leveraging interoperable data formats and automated modeling workflows, this approach supports scalable, reproducible model construction and evaluation, enabling improved efficiency and transparency in clinical decision support.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 95%
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 95%
- Real-Time Electronic Health Record Mortality Prediction During the COVID-19 Pandemic: A Prospective Cohort Study 95%
Similar papers in this journal
- A deep learning model for clinical outcome prediction using longitudinal inpatient electronic health records 95%
- Modeling physician variability to prioritize relevant medical record information 95%
- Characterizing subgroup performance of probabilistic phenotype algorithms within older adults: A case study for dementia, mild cognitive impairment, and Alzheimer’s and Parkinson’s diseases 93%
Similar papers in this journal
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 96%
- Development and Prospective Implementation of a Large Language Model based System for Early Sepsis Prediction 96%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 95%
Similar papers in this journal
- Predictability and Stability Testing to Assess Clinical Decision Instrument Performance for Children After Blunt Torso Trauma 95%
- Identification of predictive patient characteristics for assessing the probability of COVID-19 in-hospital mortality 94%
- Generalizability Challenges of Mortality Risk Prediction Models: A Retrospective Analysis on a Multi-center Database 94%
Similar papers in this journal
- Predicting Prognosis in COVID-19 Patients using Machine Learning and Readily Available Clinical Data 94%
- Image and structured data analysis for prognostication of health outcomes in patients presenting to the Emergency Department during the COVID-19 pandemic 94%
- A Deep Learning Method to Detect Opioid Prescription and Opioid Use Disorder from Electronic Health Records 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.