Fitting penalized regressions on very large genetic data using snpnet and bigstatsr
Prive, F.; Vilhjalmsson, B. J.; Aschard, H.
Show abstract
Both R packages snpnet and bigstatsr allow for fitting penalized regressions on individual-level genetic data as large as the UK Biobank. Here we benchmark bigstatsr against snpnet for fitting penalized regressions on large genetic data. We find bigstatsr to be an order of magnitude faster than snpnet when applied to the UK Biobank data (from 4.5x to 35x). We also discuss the similarities and differences between the two packages, provide theoretical insights, and make recommendations on how to fit penalized regressions in the context of genetic data.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Bayesian Hierarchical Hypothesis Testing in Large-Scale Genome-Wide Association Analysis 96%
- Estimating SNP heritability in presence of population substructure in large biobank-scale data 95%
- Using encrypted genotypes and phenotypes for collaborative genomic analyses to maintain data confidentiality 95%
Similar papers in this journal
- Using feedback in pooled experiments augmented with imputation for high genotyping accuracy at reduced cost 94%
- kGWASflow: a modular, flexible, and reproducible Snakemake workflow for k-mers-based GWAS 93%
- PyBrOpS: a Python package for breeding program simulation and optimization for multi-objective breeding 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.