Back

Metabolomic and genomic prediction of common diseases in 477,706 participants in three national biobanks

Nightingale Health Biobank Collaborative Group, ; Barrett, J. C.; Esko, T.; Fischer, K.; Jostins-Dean, L.; Jousilahti, P.; Julkunen, H.; Jaaskelainen, T.; Kerimov, N.; Kerminen, S.; Kolde, A.; Koskela, H.; Kronberg, J.; Lundgren, S. N.; Lundqvist, A.; Makela, V.; Nybo, K.; Perola, M.; Salomaa, V.; Schut, K.; Soikkeli, M.; Soininen, P.; Tiainen, M.; Tillmann`, T.; Wurtz, P.; Estonian Biobank Research Team,

2023-06-12 genetic and genomic medicine
10.1101/2023.06.09.23291213 medRxiv
Show abstract

Identifying individuals at high risk of chronic diseases via easily measured biomarkers could improve public health efforts to prevent avoidable illness and death. Here we present nuclear magnetic resonance blood metabolomics from half a million samples from three national biobanks. We built metabolomic risk scores that identify a high-risk group for each of 12 diseases that cause the most morbidity in high-income countries and show consistent cross-biobank replication of the relative risk of disease for these groups. We show that these metabolomic risk scores are more strongly associated with future disease onset than polygenic scores for most of these diseases. In a subset of 18,000 individuals with metabolomic biomarkers measured at two time points we show that people whose scores change have dramatically different future risk of disease, suggesting that repeat measurements capture the benefits of lifestyle change. We show cross-biobank calibration of our scores. Since metabolomics can be measured from a standard blood sample, we propose such tests can be feasibly implemented today in preventative health programs. One-Sentence SummaryBiomarkers from half a million blood samples identifies people at increased risk of chronic diseases and can be used for early detection today.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.