A rigorous method for integrating multiple heterogeneous databases in genetic studies
Bukszar, J.; van den Oord, E. J.
Show abstract
The large number of existing databases provides a freely available independent source of information with a considerable potential to increase the likelihood of identifying genes for complex diseases. We developed a flexible framework for integrating such heterogeneous databases into novel large scale genetic studies and implemented the methods in a freely-available, user-friendly R package called MIND. For each marker, MIND computes the posterior probability that the marker has effect in the novel data collection based on the information in all available data. MIND 1) relies on a very general model, 2) is based on the mathematical formulas that provide us with the exact value of the posterior probability, and 3) has good estimation properties because of its very efficient parameterization. For an existing data set, only the ranks of the markers are needed, where ties among the ranks are allowed. Through simulations, cross-validation analyses involving 18 GWAS, and an independent replication study of 6,544 SNPs in 6,298 samples we show that MIND 1) is accurate, 2) outperforms marker selection for follow up studies based on p-values, and 3) identifies effects that would otherwise require replication of over 20 times as many markers. AUTHOR SUMMARYThe large number of existing databases provides a freely available independent source of information with a considerable potential to increase the likelihood of identifying genes for complex diseases. We developed a flexible framework for integrating such heterogeneous databases into novel large scale genetic studies and implemented the methods in a freely-available, user-friendly R package called MIND. For each marker, MIND computes an estimate of the (posterior) probability that the marker has effect in the novel data collection based on the information in all available data. For an existing data set, only the ranks of the markers are needed to be known, where ties among the ranks are allowed. MIND 1) relies on a realistic model that takes confounding effects into account, 2) is based on the mathematical formulas that provide us with the exact value of the posterior probability, and 3) has good estimation properties because of its very efficient parameterization. Simulation, validation, and a replication study in independent samples show that MIND is accurate and greatly outperforms marker selection without using existing data sets.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Subset scanning for multi-trait analysis using GWAS summary statistics 96%
- An exact, unifying framework for region-based association testing in family-based designs, including higher criticism approaches, SKATs, multivariate and burden tests 96%
- CoMM-S2: a collaborative mixed model using summary statistics in transcriptome-wide association studies 96%
Similar papers in this journal
- Estimating the effect size of a hidden causal factor between SNPs and a continuous trait: a mediation model approach 95%
- TISSLET Tissues-based Learning Estimation for Transcriptomics 95%
- Identifying novel associations in GWAS by hierarchical Bayesian latent variable detection of differentially misclassified phenotypes 94%
Similar papers in this journal
- DOT: Gene-set analysis by combining decorrelated association statistics 96%
- Association Tests Using Copy Number Profile Curves (CONCUR) Enhances Power in Rare Copy Number Variant Analysis 96%
- Model Checking via Testing for Direct Effects in Mendelian Randomization and Transcriptome-wide Association Studies 96%
Similar papers in this journal
- kTWAS: integrating kernel-machine with transcriptome-wide association studies improves statistical power and reveals novel genes 96%
- BayesKAT: Bayesian Optimal Kernel-based Test for genetic association studies reveals joint genetic effects in complex diseases 96%
- Efficient test for deviation from Hardy Weinberg Equilibrium with known or ambiguous typing in highly polymorphic loci 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.