A Fast, Provably Accurate Approximation Algorithm for Sparse Principal Component Analysis Reveals Human Genetic Variation Across the World
Chowdhury, A.; Bose, A.; Zhou, S.; Woodruff, D. P.; Drineas, P.
Show abstract
Principal component analysis (PCA) is a widely used dimensionality reduction technique in machine learning and multivariate statistics. To improve the interpretability of PCA, various approaches to obtain sparse principal direction loadings have been proposed, which are termed Sparse Principal Component Analysis (SPCA). In this paper, we present ThreSPCA1, a provably accurate algorithm based on thresholding the Singular Value Decomposition for the SPCA problem, without imposing any restrictive assumptions on the input covariance matrix. Our thresholding algorithm is conceptually simple; much faster than current state-of-the-art; and performs well in practice. When applied to genotype data from the 1000 Genomes Project, ThreSPCA is faster than previous benchmarks, at least as accurate, and leads to a set of interpretable biomarkers, revealing genetic diversity across the world.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Private Genomes and Public SNPs: Homomorphic encryption of genotypes and phenotypes for shared quantitative genetics 96%
- Scaling the Discrete-time Wright Fisher model to biobank-scale datasets 96%
- Reflection Knockoffs via Householder Reflection: Applications in Proteomics and Genetic Fine Mapping 96%
Similar papers in this journal
Similar papers in this journal
- Survival Analysis on Rare Events Using Group-Regularized Multi-Response Cox Regression 96%
- Fast Lasso method for Large-scale and Ultrahigh-dimensional Cox Model with applications to UK Biobank 95%
- Estimating the overall fraction of phenotypic variance attributed to high-dimensional predictors measured with error 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.