Prediction of Long COVID Based on Severity of Initial COVID-19 Infection: Differences in predictive feature sets between hospitalized versus non-hospitalized index infections
Socia, D.; Larie, D.; Feuerwerker, S.; An, G.; Cockrell, C.
Show abstract
Long COVID is recognized as a significant consequence of SARS-COV2 infection. While the pathogenesis of Long COVID is still a subject of extensive investigation, there is considerable potential benefit in being able to predict which patients will develop Long COVID. We hypothesize that there would be distinct differences in the prediction of Long COVID based on the severity of the index infection, and use whether the index infection required hospitalization or not as a proxy for developing predictive models. We divide a large population of COVID patients drawn from the United States National Institutes of Health (NIH) National COVID Cohort Collaborative (N3C) Data Enclave Repository into two cohorts based on the severity of their initial COVID-19 illness and correspondingly trained two machine learning models: the Long COVID after Severe Disease Model (LCaSDM) and the Long COVID after Mild Disease Model (LCaMDM). The resulting models performed well on internal validation/testing, with a F1 score of 0.94 for the LCaSDM and 0.82 for the LCaMDM. There were distinct differences in the top 10 features used by each model, possibly reflecting the differences in type and amount of pathophysiological data between the hospitalized and non-hospitalized patients and/or reflecting different pathophysiological trajectories in the development of Long COVID. Of particular interest was the importance of Plant Hardiness Zone in the feature set for the LCaMDM, which may point to a role of climate and/or sunlight in the progression to Long COVID. Future work will involve a more detailed investigation of the potential role of climate and sunlight, as well as refinement of the predictive models as Long COVID becomes increasingly parsed into distinct clinical phenotypes.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Modeling physician variability to prioritize relevant medical record information 95%
- Trajectories: a framework for detecting temporal clinical event sequences from health data standardized to the OMOP Common Data Model 94%
- Using indication embeddings to represent patient health for drug safety studies 94%
Similar papers in this journal
- Identification of predictive patient characteristics for assessing the probability of COVID-19 in-hospital mortality 94%
- From theoretical models to practical deployment: A perspective and case study of opportunities and challenges in AI-driven healthcare research for low-income settings 94%
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 93%
Similar papers in this journal
- Mitigating Machine Learning Bias Between High Income and Low-Middle Income Countries for Enhanced Model Fairness and Generalizability 95%
- Developing Machine Learning Models for Predicting Intensive Care Unit Resource Use During the COVID-19 Pandemic 95%
- Emergency department admissions during COVID-19: explainable machine learning to characterise data drift and detect emergent health risks 95%
Similar papers in this journal
- Identification of high-risk COVID-19 patients using machine learning 94%
- A Stacked ensemble method for forecasting influenza-like illness visit volumes at emergency departments 94%
- Using mobile phone data to estimate dynamic population changes and improve the understanding of a pandemic: A case study in Andorra 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.