Deep learning for polygenic prediction: The role of heritability, interaction type and sample size
Grealey, J.; Abraham, G.; Meric, G.; Canovas, R.; Kelemen, M.; Teo, S. M.; Salim, A.; Inouye, M.; Xu, Y.
Show abstract
Polygenic scores (PGS), which aggregate the effects of genetic variants to estimate predisposition for a disease or trait, have potential clinical utility in disease prevention and precision medicine. Recently, there has been increasing interest in using deep learning (DL) methods to develop PGS, due to their strength in modelling complex non-linear relationships (such as GxG) that conventional PGS methods may not capture. However, the perceived value of DL for polygenic scores is unclear. In this study, we assess the underlying factors impacting DL performance and how they can be better utilised for PGS development. We simulate large-scale realistic genotype-to-phenotype data, with varying genetic architectures of phenotypes under quantitative control of three key components: (a) total heritability, (b) variant-variant interaction type, and (c) proportion of non-additive heritability. We compare the performance of one of most common DL methods (multi-layer perceptron, MLP) on varying training sample sizes, with two well-established PGS methods: a purely additive model (pruning and thresholding, P+T) and a machine learning method (Elastic net, EN). Our analyses show EN has consistently better overall performance across traits of different architectures and training data of different sizes. However, MLP saw the largest performance improvements as sample size increases. MLP outperformed P+T for most traits and achieves comparable performance as EN for numerous traits at the largest sample size assessed (N=100k), suggesting DL may offer some advantages in future when they can be trained on biobanks of millions of samples. We further found that one-hot encoding of variant input can improve performance of every method, particularly for traits with non-additive variance. Overall, we show how different underlying factors impact how well methods leverage non-additivity for polygenic prediction.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Incorporating family disease history and controlling case-control imbalance for population based genetic association studies 95%
- An exact, unifying framework for region-based association testing in family-based designs, including higher criticism approaches, SKATs, multivariate and burden tests 94%
- Computing Linkage Disequilibrium Aware Genome Embeddings using Autoencoders 94%
Similar papers in this journal
- Variational Autoencoder-based Model Improves Polygenic Prediction in Blood Cell Traits 95%
- Leveraging Global Genetics Resources to Enhance Polygenic Prediction Across Ancestrally Diverse Populations 94%
- A parametric bootstrap approach for computing confidence intervals for genetic correlations with application to genetically-determined protein-protein networks 93%
Similar papers in this journal
Similar papers in this journal
- Controlling for background genetic effects using polygenic scores improves the power of genome-wide association studies 95%
- Biobank-scale methods and projections for sparse polygenic prediction from machine learning 93%
- Controlling for Human Population Stratification in Rare Variant Association Studies 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.