Can longitudinal electronic health record data identify patients at higher risk of developing long COVID?
Shanmugam, P.; Bair, M.; Pendl-Robinson, E.; Hu, X. C.
Show abstract
With hundreds of millions of COVID-19 infections to date, a considerable portion of the population has developed or will develop long COVID. Understanding the prevalence, risk factors, and healthcare costs of long COVID can be of significant societal importance. To investigate the utility of large-scale electronic health record (EHR) data in identifying and predicting long COVID, we analyzed a sample of 1.23 million participants from the National COVID Cohort Collaborative (N3C), a longitudinal EHR data repository from 80 sites in the US with over 8 million COVID-19 patients. We characterized the prevalence of long COVID using a few different types of definitions to illustrate their relative strengths and weaknesses. Then we developed machine learning models to predict the risk of developing long COVID using demographic factors and comorbidity in the EHR. The risk factors for long COVID include patient age; sex; smoking status; and comorbidities characterized by the Charlson Comorbidity Index (CCI). We were able to predict three types of long COVID with low to moderate levels of accuracy (AUC 0.599 - 0.734). We found that age and CCI were most predictive of long COVID diagnosis. Ongoing work includes applying the fair machine learning framework to the long COVID predictive models. We are implementing fairness and bias mitigation methods to model fitting through the following steps, selecting fairness metrics, preparing data and model, evaluating fairness metrics, applying bias mitigation methods to the dataset, and comparing model results and fairness metrics before and after the mitigation. The objective is to achieve equalized odds, a statistical notion that ensures classification algorithms do not discriminate against protected groups (such as sex and race/ethnicity). Results from the fairness-based machine learning will be included in the conference presentation.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Image and structured data analysis for prognostication of health outcomes in patients presenting to the Emergency Department during the COVID-19 pandemic 93%
- A Deep Learning Method to Detect Opioid Prescription and Opioid Use Disorder from Electronic Health Records 93%
- Development and Evaluation of MADDIE: Method to Acquire Delivery Date Information from Electronic Health Records 91%
Similar papers in this journal
- Trajectories: a framework for detecting temporal clinical event sequences from health data standardized to the OMOP Common Data Model 94%
- Modeling physician variability to prioritize relevant medical record information 92%
- Characterizing subgroup performance of probabilistic phenotype algorithms within older adults: A case study for dementia, mild cognitive impairment, and Alzheimer’s and Parkinson’s diseases 92%
Similar papers in this journal
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 94%
- Temporally-Informed Random Forests for Suicide Risk Prediction 94%
- Use of unstructured text in prognostic clinical prediction models: a systematic review 93%
Similar papers in this journal
- COHD-COVID: Columbia Open Health Data for COVID-19 Research 95%
- Distinguishing Admissions Specifically for COVID-19 from Incidental SARS-CoV-2 Admissions: A National Retrospective EHR Study 93%
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 92%
Similar papers in this journal
- Mining for Equitable Health: Assessing the Impact of Missing Data in Electronic Health Records 94%
- ConceptWAS: a high-throughput method for early identification of COVID-19 presenting symptoms 93%
- A Deep Learning Approach for Transgender and Gender Diverse Patient Identification in Electronic Health Records 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.