Back

Quality control and removal of technical variation of NMR metabolic biomarker data in ~120,000 UK Biobank participants

Ritchie, S. C.; Surendran, P.; Karthikeyan, S.; Lambert, S. A.; Bolton, T.; Pennells, L.; Danesh, J.; Di Angelantonio, E.; Butterworth, A. S.; Inouye, M.

2021-09-27 cardiovascular medicine
10.1101/2021.09.24.21264079 medRxiv
Show abstract

Metabolic biomarker data quantified by nuclear magnetic resonance (NMR) spectroscopy has recently become available in UK Biobank. Here, we describe procedures for quality control and removal of technical variation for this biomarker data, comprising 249 circulating metabolites, lipids, and lipoprotein sub-fractions on approximately 121,000 participants. We identify and characterise technical and biological factors associated with individual biomarkers and find that linear effects on individual biomarkers can combine in a non-linear fashion for 61 composite biomarkers and 81 biomarker ratios. We create an R package, ukbnmr, for extracting and normalising the metabolic biomarker data, then use ukbnmr to remove unwanted variation from the UK Biobank data. We make available code for re-deriving the 61 composite biomarkers and 81 ratios, and for further derivation of 76 additional biomarker ratios of potential biological significance. Finally, we demonstrate that removal of technical variation leads to increased signal for genetic and epidemiological studies of the NMR metabolic biomarkers in UK Biobank.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.