Comorbidity Exposure-Window Definitions and Multidimensional Disparities in Long COVID Risk: Evidence from a U.S. National Cohort (2020-2024)
Chen, Y.; Chen, Z.; Yang, G.; Li, B.; Ogunyemi, K. O.; Liu, J.; Luo, F.; Ke, Y.; Martinez, L.; Chen, X.; Rajbhandari, J.; Shen, Y.
Show abstract
Long COVID (LC) affects millions of individuals worldwide, particularly those with preexisting comorbidities. However, whether these comorbidities should be defined before SARS-CoV-2 infection or before LC diagnosis remains unresolved, and this methodological choice may substantially bias estimates of comorbidity-associated LC risk. In addition, most previous studies were conducted during earlier phases of the pandemic and relied on relatively small or geographically restricted cohorts, limiting understanding of temporal trends and population disparities in LC risk. Leveraging Electronic Health Records (EHR) from 6,130,413 adults with documented COVID-19 across 49 U.S. states in the National COVID Cohort Collaborative (N3C) from 2020 to 2024, we evaluated the impact of different comorbidity exposure-window definitions on LC risk estimation. We utilized ensemble cross-fitted double/debiased machine learning to adjust for complex individual- and county-level confounders. Across the 16 major comorbidities evaluated, defining conditions before SARS-CoV-2 infection, rather than before LC diagnosis, yielded 23%--115% higher adjusted attributable risks and 6%--37% higher adjusted relative risks. Additionally, comorbidity-associated risks generally declined from 2020 to 2024, with substantial demographic, socioeconomic, geographic, and multimorbidity-related disparities persisting throughout the study period. These findings identify temporal exposure-window specification as a major source of bias in LC epidemiology. Failure to distinguish preexisting comorbidities from conditions identified during postinfection follow-up can substantially alter estimates of disease burden, the identification of high-risk populations, and the interpretation of temporal and geographic disparities. More broadly, our results highlight how temporal misclassification of exposures in longitudinal EHR studies can distort risk attribution and population-level inference.
Matching journals
The top 12 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Risk factors for long COVID: analyses of 10 longitudinal studies and electronic health records in the UK 94%
- Trends and associated factors for Covid-19 hospitalisation and fatality risk in 2.3 million adults in England 93%
- A unified framework for estimating country-specific cumulative incidence for 18 diseases stratified by polygenic risk 92%
Similar papers in this journal
- Sociodemographic Characteristics and Longitudinal Progression of Multimorbidity: A Multistate Modelling Analysis of a Large Primary Care Records Dataset in England 92%
- Overall and cause-specific hospitalisation and death after COVID-19 hospitalisation in England: cohort study in OpenSAFELY using linked primary care, secondary care and death registration data 92%
- Rapid Epidemiological Analysis of Comorbidities and Treatments as risk factors for COVID-19 in Scotland (REACT-SCOT): a population-based case-control study 91%
Similar papers in this journal
Similar papers in this journal
- A systematic analysis of the contribution of genetics to multimorbidity and comparisons with primary care data 94%
- Refinement of post-COVID condition core symptoms, subtypes, determinants, and health impacts: A cohort study integrating real-world data and patient-reported outcomes 93%
- Older biological age is associated with adverse COVID-19 outcomes: A cohort study in UK Biobank 92%
Similar papers in this journal
- Estimating the Burden of Influenza on Daily Activity at Population Scale Using Commercial Wearable Sensors 92%
- Effectiveness of the Single-Dose Ad26.COV2.S COVID Vaccine 91%
- Multimorbidity and risk of incident dementia: role of disease clusters and genetic risk for dementia in a cohort of 206,960 participants 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.