Who has long-COVID? A big data approach
Pfaff, E. R.; Girvin, A. T.; Bennett, T. D.; Bhatia, A.; Brooks, I. M.; Deer, R. R.; Dekermanjian, J. P.; Jolley, S. E.; Kahn, M. G.; McMurry, J. A.; Moffitt, R. A.; Walden, A.; Chute, C. G.; Haendel, M. A.; The N3C Consortium,
Show abstract
BackgroundPost-acute sequelae of SARS-CoV-2 infection (PASC), otherwise known as long-COVID, have severely impacted recovery from the pandemic for patients and society alike. This new disease is characterized by evolving, heterogeneous symptoms, making it challenging to derive an unambiguous long-COVID definition. Electronic health record (EHR) studies are a critical element of the NIH Researching COVID to Enhance Recovery (RECOVER) Initiative, which is addressing the urgent need to understand PASC, accurately identify who has PASC, and identify treatments. MethodsUsing the National COVID Cohort Collaboratives (N3C) EHR repository, we developed XGBoost machine learning (ML) models to identify potential long-COVID patients. We examined demographics, healthcare utilization, diagnoses, and medications for 97,995 adult COVID-19 patients. We used these features and 597 long-COVID clinic patients to train three ML models to identify potential long-COVID patients among (1) all COVID-19 patients, (2) patients hospitalized with COVID-19, and (3) patients who had COVID-19 but were not hospitalized. FindingsOur models identified potential long-COVID patients with high accuracy, achieving areas under the receiver operator characteristic curve of 0.91 (all patients), 0.90 (hospitalized); and 0.85 (non-hospitalized). Important features include rate of healthcare utilization, patient age, dyspnea, and other diagnosis and medication information available within the EHR. Applying the "all patients" model to the larger N3C cohort identified 100,263 potential long-COVID patients. InterpretationPatients flagged by our models can be interpreted as "patients likely to be referred to or seek care at a long-COVID specialty clinic," an essential proxy for long-COVID diagnosis in the current absence of a definition. We also achieve the urgent goal of identifying potential long-COVID patients for clinical trials. As more data sources are identified, the models can be retrained and tuned based on study needs. FundingThis study was funded by NCATS and NIH through the RECOVER Initiative.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Continuous-Time and Dynamic Suicide Attempt Risk Prediction with Neural Ordinary Differential Equations 95%
- Finding Long-COVID: Temporal Topic Modeling of Electronic Health Records from the N3C and RECOVER Programs 95%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 95%
Similar papers in this journal
Similar papers in this journal
- Real-Time Electronic Health Record Mortality Prediction During the COVID-19 Pandemic: A Prospective Cohort Study 95%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 95%
- Validation of a Derived International Patient Severity Algorithm to Support COVID-19 Analytics from Electronic Health Record Data 95%
Similar papers in this journal
- A machine learning-based phenotype for long COVID in children: an EHR-based study from the RECOVER program 96%
- Development of a prediction model for 30-day COVID-19 hospitalization and death in a national cohort of Veterans Health Administration patients – March 2022 - April 2023 94%
- Clinical prediction rule for SARS-CoV-2 infection from 116 U.S. emergency departments 94%
Similar papers in this journal
- Automated Interpretable Discovery of Heterogeneous Treatment Effectiveness: A Covid-19 Case Study 94%
- Natural language processing for scalable feature engineering and ultra-high-dimensional confounding adjustment in healthcare database studies 94%
- ConceptWAS: a high-throughput method for early identification of COVID-19 presenting symptoms 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.