GARSA: An integrative pipeline for genome wide association studies and polygenic risk score inference in admixed human populations
Rossi, F.; Patane, J.; de Souza, V.; Neyra, J.; Rosa, R.; Krieger, J.; Teixeira, S.
Show abstract
Genome-wide association studies (GWAS) and polygenic risk scores (PRS) are multistep analytical tools to identify genetic variants and to assess their contribution to phenotypes/diseases. These analyses are evolving and becoming instrumental to understand the genetic architecture of complex phenotypes/diseases. Nevertheless, to date, there is no single solution incorporating all major steps related to those analyses combined with robust populational bias correction. Here, we describe a semi-automated pipeline unifying steps involved in GWAS and PRS including widely used software. Our pipeline handles quality control (QC), GWAS, and PRS steps, managing different types of input/output files. Furthermore, it includes robust bias correction steps, such as inference of kinship matrix with correction for population structure, use of principal component analysis (PCA) with detection and removal of outlier variant followed by re-projection of related individuals (if desired), generation of PCA figures that assist in setting the best number of principal components (PCs) for association analysis, availability of mixed models, use of recommended software for GWAS based on population size, and a Markov chain Monte Carlo (MCMC) method to estimate best set of PRS parameters. Finally, we tested GARSA pipeline in a family-based Brazilian admixed population and demonstrated that the corrections implemented indeed mitigate bias in downstream analysis. The pipeline can be implemented on personal or server-side environments. AvailabilityThe development version (open-source) is available in https://github.com/LGCM-OpenSource/GARSA ContactFernando P. N. Rossi - fernando.rossi@hc.fm.usp.br; Jose S. L. Patane - jose.patane@hc.fm.usp.br Supplementary informationSupplementary tutorial.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- MetaPhat: Detecting and decomposing multivariate associations from univariate genome-wide association statistics 94%
- MicroHapDB: a portable and extensible database of all published microhaplotype marker and frequency data 93%
- SECNVs: A Simulator of Copy Number Variants and Whole-Exome Sequences from Reference Genomes 93%
Similar papers in this journal
- VarGenius-HZD allows accurate detection of rare homozygous or hemizygous deletions in targeted sequencing leveraging breadth of coverage 92%
- A novel framework for analysis of the shared genetic background of correlated traits 91%
- The FORCE panel: An all-in-one SNP marker set for confirming investigative genetic genealogy leads and for general forensic applications 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.