Back

Integration of DNA methylation datasets for individual prediction

Merzbacher, C.; Ryan, B.; Goldsborough, T.; Hillary, R. F.; Campbell, A.; Murphy, L.; McIntosh, A. M.; Liewald, D.; Harris, S. E.; McRae, A. F.; Cox, S. R.; Cannings, T. I.; Vallejos, C. A.; McCartney, D. L.; Marioni, R. E.

2023-03-22 genetic and genomic medicine
10.1101/2023.03.22.23287572 medRxiv
Show abstract

BackgroundEpigenetic scores (EpiScores) can provide blood-based biomarkers of lifestyle and disease risk. Projecting a new individual onto a reference panel would aid precision medicine and risk communication but is challenging due to the separation of technical and biological sources of variation with array data. Normalisation methods can standardize data distributions but may also remove population-level biological variation. MethodsWe compared two independent birth cohorts (Lothian Birth Cohorts of 1921 and 1936 - nLBC1921 = 387 and nLBC1936 = 498) with DNA methylation assessed at the same chronological age (79 years) and processed in the same lab but in different years and experimental batches. We examined the effect of 15 normalisation methods on a BMI EpiScore (trained in an external cohort of 18,413 individuals) when the cohorts were normalised separately and together. ResultsThe BMI EpiScore explained a maximum variance of R2=24.5% in BMI in LBC1936 after SWAN normalisation. Although there were differences in the variance explained across cohorts, the normalisation methods made minimal differences to the estimates within cohorts. Conversely, a range of absolute differences were seen for individual-level EpiScore estimates when cohorts were normalised separately versus together. While within-array methods result in identical BMI EpiScores whether a cohort was normalised on its own or together with the second dataset, a range of differences were observed for between-array methods. ConclusionsUsing normalisation methods that give similar EpiScores whether cohorts are analysed separately or together will minimise technical variation when projecting new data onto a reference panel. These methods are especially important for cases where when raw data and joint normalisation of cohorts is not possible or is computationally expensive.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.