Improving Neural Networks for Genotype-Phenotype Prediction Using Published Summary Statistics
Cui, T.; El Mekkaoui, K.; Havulinna, A. S.; Marttinen, P.; Kaski, S.
Show abstract
Phenotype prediction is a necessity in numerous applications in genetics. However, when the size of the individual-level data of the cohort of interest is small, statistical learning algorithms, from linear regression to neural networks, usually fail due to insufficient data. Fortunately, summary statistics from genome-wide association studies (GWAS) on other large cohorts are often publicly available. In this work, we propose a new regularization method, namely, main effect prior (MEP), for making use of GWAS summary statistics from external datasets. The main effect prior is generally applicable for machine learning algorithms, such as neural networks and linear regression. With simulation and real-world experiments, we show empirically that MEP improves the prediction performance on both homogeneous and heterogeneous datasets. Moreover, deep neural networks with MEP outperform standard baselines even when the training set is small.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Sparse Multitask group Lasso for Genome-Wide Association Studies 96%
- RCFGL: Rapid Condition adaptive Fused Graphical Lasso and application to modeling brain region co-expression networks 96%
- Biological networks and GWAS: comparing and combining network methods to understand the genetics of familial breast cancer susceptibility in the GENESIS study 95%
Similar papers in this journal
Similar papers in this journal
- Identification of Significant Gene Expression Changes in Multiple Perturbation Experiments using Knockoffs 96%
- kTWAS: integrating kernel-machine with transcriptome-wide association studies improves statistical power and reveals novel genes 96%
- CoxMDS: Multiple Data Splitting for High-dimensional Mediation Analysis with Survival Outcomes in Epigenome-wide Studies 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.