Tractor Workflow Pipeline: A Scalable Nextflow Framework for Local Ancestry-Aware Genome-Wide Association Studies
Shah, N. N.; Tan, T.; Honorato-Mauer, J.; Lin, Y.-S.; Maihofer, A. X.; Zai, C. C.; PGC-PTSD Ancestry Working Group, ; Santoro, M. L.; Nievergelt, C. M.; Atkinson, E. G.
Show abstract
The routine exclusion of admixed individuals from traditional Genome-Wide Association Studies (GWAS) due to concerns about spurious associations has hindered genetic analyses involving multiple ancestries. Tractor GWAS addresses this issue by incorporating local ancestry into its analysis, empowering identification of ancestry-enriched hits and generating ancestry-specific summary statistics. However, Tractor requires accurate genomic phasing and local ancestry inference as prerequisite steps, which requires additional bioinformatics expertise and decision points regarding reference panel setup. To streamline, harmonize, and automate this process, we present a scalable Nextflow workflow that integrates all necessary steps, minimizing the need for manual intervention while remaining modular and customizable. The workflow supports multiple commonly used tools and offers flexibility in how Tractor is implemented. To demonstrate its utility, we applied this pipeline to analyze 32 blood biomarkers in 6,245 two-way AFR-EUR admixed individuals from the UK Biobank. This pipeline ran efficiently at scale, replicated known associations, and identified novel ancestry-specific loci. These novel associations were largely driven by variants present on African ancestral tracts but absent from European tracts, underscoring the value of local ancestry-aware methods in uncovering previously missed genetic signals. By enabling the efficient analysis of admixed individuals, our workflow facilitates Tractor use, paving the way for more broader genetic discovery.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CLUES2 Companion: Computational pipelines to estimate, visualize, and date selection on multi-locus sites 96%
- ntRoot: Computational Inference of Human Ancestry at Scale from Genomic Data 94%
- AnnSQL: A Python SQL-based package for fast large-scale single-cell genomics analysis using minimal computational resources 94%
Similar papers in this journal
- GenoTools: An Open-Source Python Package for Efficient Genotype Data Quality Control and Analysis 95%
- kGWASflow: a modular, flexible, and reproducible Snakemake workflow for k-mers-based GWAS 95%
- Low-pass sequencing plus imputation using avidity sequencing displays comparable imputation accuracy to sequencing by synthesis while reducing duplicates 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.