Proteomic prediction of common and rare diseases
Carrasco-Zanini, J.; Pietzner, M.; Davitte, J.; Surendran, P.; Croteau-Chonka, D. C.; Robins, C.; Torralbo, A.; Tomlinson, C.; Fitzpatrick, N.; Ytsma, C.; Kanno, T.; Gade, S.; Freitag, D.; Ziebell, F.; Denaxas, S.; Betts, J. C.; Wareham, N. J.; Hemingway, H.; Scott, R. A.; Langenberg, C.
Show abstract
BackgroundFor many diseases there are delays in diagnosis due to a lack of objective biomarkers for disease onset. Whether measuring thousands of proteins offers predictive information across a wide range of diseases is unknown. MethodsIn 41,931 individuals from the UK Biobank Pharma Proteomics Project (UKB-PPP), we integrated [~]3000 plasma proteins with clinical information to derive sparse prediction models for the 10-year incidence of 218 common and rare diseases (81 - 6038 cases). We compared prediction models based on proteins with a) basic clinical information alone, b) basic clinical information + 37 clinical biomarkers, and c) genome-wide polygenic risk scores. ResultsFor 67 pathologically diverse diseases, a model including as few as 5 to 20 proteins was superior to clinical models (median delta C-index = 0.07; range = 0.02 - 0.31) and to clinical models with biomarkers for 52 diseases. In multiple myeloma, for example, a set of 5 proteins significantly improved prediction over basic clinical information (delta C-index = 0.25 (95% confidence interval 0.20 - 0.29)). At a 5% false positive rate (FPR), proteomic prediction (5 proteins) identified individuals at high risk of multiple myeloma (detection rate (DR) = 50%), non-Hodgkin lymphoma (DR = 55%) and motor neuron disease (DR = 29%). At a 20% FPR, proteomic prediction identified individuals at high-risk for pulmonary fibrosis (DR= 80%) and dilated cardiomyopathy (DR = 75%). ConclusionsSparse plasma protein signatures offer novel, clinically useful prediction of common and rare diseases, through disease-specific proteins and protein predictors shared across multiple diseases. (Funded by Medical Research Council, NIHR, Wellcome Trust.)
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- The Interpretable Multimodal Machine Learning (IMML) framework reveals pathological signatures of distal sensorimotor polyneuropathy 95%
- Deep Proteome Profiling of Metabolic Dysfunction-Associated Steatotic Liver Disease 94%
- Complex patterns of multimorbidity associated with severe COVID-19 and Long COVID 93%
Similar papers in this journal
Similar papers in this journal
- Machine learning-guided deconvolution of plasma protein levels 94%
- Multi-cohort, cross-species urinary proteomics reveals signatures of LRRK2 dysfunction in Parkinsons disease 93%
- A tissue-aware machine learning framework enhances the mechanistic understanding and genetic diagnosis of Mendelian and rare diseases 92%
Similar papers in this journal
- A Neanderthal OAS1 isoform Protects Against COVID-19 Susceptibility and Severity: Results from Mendelian Randomization and Case-Control Studies 93%
- Actionable druggable genome-wide Mendelian randomization identifies repurposing opportunities for COVID-19 93%
- Genome-wide polygenic score with APOL1 risk genotypes predicts chronic kidney disease across major continental ancestries 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.