Predicting Long COVID in the National COVID Cohort Collaborative Using Super Learner
Butzin-Dozier, Z.; Ji, Y.; Li, H.; Coyle, J.; Shi, J.; Phillips, R. V.; Mertens, A.; Pirracchio, R.; van der Laan, M. J.; Patel, R. C.; Colford, J. M.; Hubbard, A. E.; on behalf of the National COVID Cohort Collaborative (N3C) Consortium*,
Show abstract
Post-acute Sequelae of COVID-19 (PASC), also known as Long COVID, is a broad grouping of a range of long-term symptoms following acute COVID-19 infection. An understanding of characteristics that are predictive of future PASC is valuable, as this can inform the identification of high-risk individuals and future preventative efforts. However, current knowledge regarding PASC risk factors is limited. Using a sample of 55,257 participants from the National COVID Cohort Collaborative, as part of the NIH Long COVID Computational Challenge, we sought to predict individual risk of PASC diagnosis from a curated set of clinically informed covariates. We predicted individual PASC status, given covariate information, using Super Learner (an ensemble machine learning algorithm also known as stacking) to learn the optimal, AUC-maximizing combination of gradient boosting and random forest algorithms. We were able to predict individual PASC diagnoses accurately (AUC 0.947). Temporally, we found that baseline characteristics were most predictive of future PASC diagnosis, compared with characteristics immediately before, during, or after COVID-19 infection. This finding supports the hypothesis that clinicians may be able to accurately assess the risk of PASC in patients prior to acute COVID diagnosis, which could improve early interventions and preventive care. We found that medical utilization, demographics, anthropometry, and respiratory factors were most predictive of PASC diagnosis. This highlights the importance of respiratory characteristics in PASC risk assessment. The methods outlined here provide an open-source, applied example of using Super Learner to predict PASC status using electronic health record data, which can be replicated across a variety of settings.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Performance of Existing and Novel Symptom- and Antigen Testing-Based COVID-19 Case Definitions in a Community Setting 94%
- A hypothetical intervention to reduce inequities in anxiety for Multiracial people: simulating an intervention on childhood adversity 93%
- Collaborative Cohort of Cohorts for COVID-19 Research (C4R) Study: Study Design 92%
Similar papers in this journal
- Development of a prediction model for 30-day COVID-19 hospitalization and death in a national cohort of Veterans Health Administration patients – March 2022 - April 2023 95%
- A machine learning-based phenotype for long COVID in children: an EHR-based study from the RECOVER program 94%
- Diagnostic Value of Symptoms for Pediatric SARS-CoV-2 Infection in a Primary Care Setting 94%
Similar papers in this journal
Similar papers in this journal
- Modeling physician variability to prioritize relevant medical record information 92%
- Evaluation of a Machine Learning Approach Utilizing Wearable Data for Prediction of SARS-CoV-2 Infection in Healthcare Workers 91%
- Clinical Study Applying Machine Learning to Detect a Rare Disease: Results and Lessons Learned 91%
Similar papers in this journal
- Generalizability Challenges of Mortality Risk Prediction Models: A Retrospective Analysis on a Multi-center Database 94%
- Geographical validation of the Smart Triage Model by age group 92%
- Predictability and Stability Testing to Assess Clinical Decision Instrument Performance for Children After Blunt Torso Trauma 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.