Clustering of patient comorbidities within electronic medical records enables high-precision COVID-19 mortality prediction
Le Lannou, E.; Post, B.; Haar, S.; Brett, S.; Kadirvelu, B.; Faisal, A. A.
Show abstract
We present an explainable AI framework to predict mortality after a positive COVID-19 diagnosis based solely on data routinely collected in electronic healthcare records (EHRs) obtained prior to diagnosis. We grounded our analysis on the [1/2] Million people UK Biobank and linked NHS COVID-19 records. We developed a method to capture the complexities and large variety of clinical codes present in EHRs, and we show that these have a larger impact on risk than all other patient data but age. We use a form of clustering for natural language processing of the clinical codes, specifically, topic modelling by Latent Dirichlet Allocation (LDA), to generate a succinct digital fingerprint of a patients full secondary care clinical history, i.e. their comorbidities and past interventions. These digital comorbidity fingerprints offer immediately interpretable clinical descriptions that are meaningful, e.g. grouping cardiovascular disorders with common risk factors but also novel groupings that are not obvious. The comorbidity fingerprints differ in both their breadth and depth from existing observational disease associations in the COVID-19 literature. Taking this data-driven approach allows us to avoid human-induction bias and confirmation bias during selection of what are important potential predictors of COVID-19 mortality. Together with age, these digital fingerprints are the single most important factor in our predictor. This holds the potential for improving individual risk profiling for clinical decisions and the identification of groups for public health interventions such as vaccine programmes. Combining our digital precondition fingerprints with demographic characteristics allow us to match or exceed the performance of existing state-of-the-art COVID-19 mortality predictors (EHCF) which have been developed through expert consensus. Our precondition fingerprinting and entire mortality prediction analytics pipeline are designed so as to be rapidly redeployable, e.g. for COVID-19 variants or other pre-existing diseases.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Machine learning approach to dynamic risk modeling of mortality in COVID-19: a UK Biobank study 98%
- Developing And Validating COVID-19 Adverse Outcome Risk Prediction Models From A Bi-National European Cohort Of 5594 Patients 96%
- Personalized survival probabilities for SARS-CoV-2 positive patients by explainable machine learning 95%
Similar papers in this journal
- An external validation of the QCovid risk prediction algorithm for risk of mortality from COVID-19 in adults: national validation cohort study in England 95%
- Understanding COVID-19 trajectories from a nationwide linked electronic health record cohort of 56 million people: phenotypes, severity, waves & vaccination 95%
- Predicting hospital-onset COVID-19 infections using dynamic networks of patient contacts: an observational study 94%
Similar papers in this journal
- Pretrained Patient Trajectories for Adverse Drug Event Prediction Using Common Data Model-based Electronic Health Records 93%
- The Interpretable Multimodal Machine Learning (IMML) framework reveals pathological signatures of distal sensorimotor polyneuropathy 92%
- Sex-specific transcriptome similarity networks elucidate comorbidity relationships 91%
Similar papers in this journal
- Identifying markers of health-seeking behaviour and healthcare access in UK electronic health records 94%
- Modifiable and non-modifiable risk factors for COVID-19: results from UK Biobank 93%
- Development and validation of an algorithm to estimate the risk of severe complications of COVID-19 to prioritise vaccination 92%
Similar papers in this journal
- The adverse impact of COVID-19 pandemic on cardiovascular disease prevention and management in England, Scotland and Wales: A population-scale analysis of trends in medication data 93%
- Evaluating and Mitigating Limitations of Large Language Models in Clinical Decision Making 91%
- No statistical evidence for an effect of CCR5-Δ32 on lifespan in the UK Biobank cohort 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.