Haplomics: A Snakemake pipeline for haplotype-based association analysis in multi-omics studies
Ghasemi, D.; Foco, L.; Fujii, R.; Rainer, J.; Feli, M.; Ralser, M.; Domingues, F. S.; Pramstaller, P. P.; Pattaro, C.
Show abstract
SummaryGenome-wide association studies (GWAS) generated thousands of loci associated with complex traits and diseases. However, to characterize of the pleiotropic, molecular and population genetic bounds of uncovered loci, investigations are conducted, that remain conceptually and practically separated. Also in the best case, such investigations proceed by one locus at a time. To address these limitations, we introduce an efficient and reproducible Snakemake pipeline for executing haplotype-based association analysis on GWAS-identified genetic loci, which is especially helpful in samples enriched with molecular omics data. Haplomics takes as input the genetic coordinates of each locus along with all available clinical and molecular phenotypes, the necessary covariates, and VCF genotype files, to reconstruct haplotypes and test them for associations with the phenotypes. The reconstructed haplotypes, the annotation of included variants, and association results are graphically displayed in an HTML report. We tested Haplomics in population-based study sample encompassing 391 traits, including 72 clinical markers, 171 serum metabolites, 148 plasma protein concentrations, and whole-exome sequencing (WES) imputed genotypes. We estimated WES-based haplotypes at 11 kidney function genetic loci from a GWAS and conducted association analyses throughout, identifying 19 significant associations after multiple testing correction. Haplomics is a scalable, easy-to-use and fast haplotype reconstruction and association pipeline that makes it possible to jointly conduct molecular and population-genetic characterization of multiple GWAS loci in unified analysis framework. Availability and implementationHaplomics is freely available on GitHub at https://github.com/dariushghasemi/haplomics.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- VAREANT : a bioinformatics application for gene variant reduction and annotation 94%
- RegionScan: A comprehensive R package for region-level genome-wide association testing with integration and visualization of multiple-variant and single-variant hypothesis testing 93%
- RNApysoforms: Fast rendering interactive visualization of RNA isoform structure and expression in Python 92%
Similar papers in this journal
- Somalier: rapid relatedness estimation for cancer and germline studies using efficient genome sketches 91%
- Nanopore sequencing with unique molecular identifiers enables accurate mutation analysis and haplotyping in the complex Lipoprotein(a) KIV-2 VNTR 91%
- Validation of a Trans-Ancestry Polygenic Risk Score for Type 2 Diabetes in Diverse Populations 90%
Similar papers in this journal
- bulkAnalyseR: An accessible, interactive pipeline for analysing and sharing bulk multi-modal sequencing data 94%
- kTWAS: integrating kernel-machine with transcriptome-wide association studies improves statistical power and reveals novel genes 93%
- A novel haplotype-based eQTL approach identifies genetic associations not detected through conventional SNP-based methods 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.