Predicting COVID-19 hospitalisation and common disease risk from comorbid diagnoses in 13 million individuals
Liu, H.; Mizani, M. A.; Zhao, Y.; Wood, A.; Inouye, M.; Price, A. L.; Jiang, X.; CVD-COVID-UK/COVID-IMPACT Consortium,
Show abstract
Predicting disease risk from prior diagnoses is fundamental to clinical decision-making, particularly during health emergencies such as the COVID-19 pandemic, when individuals with long-term conditions may be disproportionately vulnerable to adverse outcomes. Despite intense interest in developing models to predict disease risk from prior diagnoses (1-3), most prediction models do not estimate effects of each prior diagnosis on disease risk conditional on other diagnoses, limiting interpretability and clinical utility. We developed the Comorbidity Risk Score (CRS), trained on 13 million individuals (age 40-69) from linked electronic health record (EHR) datasets of the entire population of England, to predict COVID-19 hospitalisation and 87 other disease outcomes. CRS was trained at close to saturated sample size and precisely estimated the effects of 212 prior diagnoses on the 88 disease outcomes, conditional on all other prior diagnoses. Correlations of CRS effect sizes across outcomes (e.g. 0.76 for myocardial infarction vs. hyperlipidaemia) matched the corresponding genetic correlations (e.g. 0.79 for myocardial infarction vs. hyperlipidaemia), confirming that comorbidity architectures capture disease aetiology. On average, CRS identified 5% of the population with 3.4-fold higher disease risk, including myocardial infarction (4.4-fold), lung cancer (6.5-fold), and COVID-19 hospitalisation (6.3-fold). Using prior diagnoses alone, CRS outperformed state-of-the-art clinical COVID-19 models (4). Furthermore, CRS (N=13 million) substantially outperformed state-of-the-art AI (1) (N=0.5 million) and linear (3) (N=0.5 million) models in predicting disease risk, suggesting that training sample size outweighs model complexity. CRS attained near-perfect transferability across self-reported ethnicities (e.g., Black vs. White: AUROC ratio = 97.3%). Finally, CRS distinguished independently predictive comorbidities from indirect associations, e.g., lipid metabolism disorder was a strong predictor of myocardial infarction risk but not ischaemic stroke, after conditioning on other prior diagnoses. In conclusion, CRS provides a comprehensive resource for understanding the impact of comorbidities on COVID-19 and other future diseases, revealing disease aetiology while enabling powerful prediction of disease risk.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Cohort Design and Natural Language Processing to Reduce Bias in Electronic Health Records Research: The Community Care Cohort Project 93%
- Identifying clusters of people with Multiple Long-Term Conditions using Large Language Models: a population-based study 93%
- Novel clinical subphenotypes in COVID-19: derivation, validation, prediction, temporal patterns, and interaction with social determinants of health 92%
Similar papers in this journal
Similar papers in this journal
- Complex patterns of multimorbidity associated with severe COVID-19 and Long COVID 92%
- Pretrained Patient Trajectories for Adverse Drug Event Prediction Using Common Data Model-based Electronic Health Records 90%
- Genetic, demographic and clinical variables act synergistically to impact neurodevelopmental outcomes in children with single ventricle heart disease 90%
Similar papers in this journal
- A systematic analysis of the contribution of genetics to multimorbidity and comparisons with primary care data 92%
- Multi-ancestry omic Mendelian randomization revealing putative drug targets of COVID-19 severity 91%
- Distinct metabolic features of genetic liability to type 2 diabetes and coronary artery disease: a reverse Mendelian randomization study 89%
Similar papers in this journal
- Polygenic score informed by genome-wide association studies of multiple ancestries and related traits improves risk prediction for coronary artery disease 92%
- Genome-wide polygenic score with APOL1 risk genotypes predicts chronic kidney disease across major continental ancestries 92%
- Actionable druggable genome-wide Mendelian randomization identifies repurposing opportunities for COVID-19 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.