Back

SAGA (Simplified Association Genomewide Analyses): a user-friendly Pipeline to Democratize Genome-Wide Association Studies

Cieza, B.; Pandey, N.; Ruhela, V.; Ali, S.; tosto, g.

2025-08-29 bioinformatics
10.1101/2025.08.25.672146 bioRxiv
Show abstract

Genome-wide association studies (GWAS) have enabled clinicians and researchers to identify genetic variants linked to complex traits and diseases(1-3). However, GWAS still face several challenges, particularly regarding accessibility and reproducibility (4-6). Conducting these analyses often requires substantial bioinformatics expertise for data preprocessing, software installation, and scripting(7-10). We then developed SAGA ("Simplified Association Genome-wide Analyses"), a BASH-based, open-source, fully automated pipeline that integrates three widely adopted tools--PLINK(11), GMMAT(12), and SAIGE(13)--for accessible, robust, and reproducible GWAS. After installation, users simply need to provide genotype and phenotype files in standard formats. The pipeline automates preprocessing, association testing, and visualization, outputting summary statistics, Manhattan plots, and quantile-quantile plots. SAGA enables robust GWAS for users without scripting experience, expanding access to complex genetic analyses.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.