PCA outperforms popular hidden variable inference methods for QTL mapping
Zhou, H. J.; Li, L.; Li, Y.; Li, W.; Li, J. J.
Show abstract
Estimating and accounting for hidden variables is widely practiced as an important step in molecular quantitative trait locus (molecular QTL, henceforth "QTL") analysis for improving the power of QTL identification. However, few benchmark studies have been performed to evaluate the efficacy of the various methods developed for this purpose. Here we benchmark popular hidden variable inference methods including surrogate variable analysis (SVA), probabilistic estimation of expression residuals (PEER), and hidden covariates with prior (HCP) against principal component analysis (PCA)--a well-established dimension reduction and factor discovery method--via 362 synthetic and 110 real data sets. We show that PCA not only underlies the statistical methodology behind the popular methods but is also orders of magnitude faster, better-performing, and much easier to interpret and use. To help researchers use PCA in their QTL analysis, we provide an R package PCAForQTL along with a detailed guide, both of which are freely available at https://github.com/heatherjzhou/PCAForQTL.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Primo: integration of multiple GWAS and omics QTL summary statistics for elucidation of molecular mechanisms of trait-associated SNPs and detection of pleiotropy in complex traits 97%
- ClipperQTL: ultrafast and powerful eGene identification method 96%
- scDesign2: a transparent simulator that generates high-fidelity single-cell gene expression count data with gene correlations captured 96%
Similar papers in this journal
Similar papers in this journal
- WEVar: a novel statistical learning framework for predicting noncoding regulatory variants 95%
- kTWAS: integrating kernel-machine with transcriptome-wide association studies improves statistical power and reveals novel genes 95%
- Comparison of sparse biclustering algorithms for gene expression datasets 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.