Back

IMPALA: A Comprehensive Pipeline for Detecting and Elucidating Mechanisms of Allele Specific Expression in Cancer

Chang, G.; Porter, V.; O'Neill, K.; Akbari, V.; Culibrk, L.; Marco, M. A.; Jones, S.

2023-09-12 bioinformatics
10.1101/2023.09.11.555771 bioRxiv
Show abstract

SummaryAllele-specific expression (ASE), where transcripts from one allele are more abundant than transcripts from the other, can arise from various genetic mechanisms and has implications for gene regulation and disease. We present IMPALA (Integrated Mapping and Profiling of Allelically-expressed Loci with Annotations), a versioned and containerized pipeline for detecting ASE in samples including cancer genomes. IMPALA leverages RNA sequencing data and, optionally, phased variant, copy number variant (CNV), allelic methylation, and mutation data to identify ASE genes and uncover underlying regulatory mechanisms. IMPALA incorporates the MBASED framework for ASE detection, and outputs a comprehensive summary table and informative figures to visualize the genomic distribution of ASE genes and their correlation with potential regulatory causes. We applied IMPALA to a cancer sample and identified thousands of genes with ASE and highlighted potential somatic events that may have influenced ASE of these genes. ASE data can be used to detect the downstream consequences of genomic alterations, which facilitates the identification of dysregulated cancer-related genes. IMPALA thus provides researchers with a powerful tool for both ASE analysis and for investigating genetic factors correlated with ASE. Availability and implementationIMPALA is licensed under GNU General Public License v3.0 and freely available at https://github.com/bcgsc/IMPALA and https://doi.org/10.5281/zenodo.8019168 with documentation and tutorial. Contactsjones@bcgsc.ca Supplemental informationSupplemental materials are available at Bioinformatics online. Issue section: Gene expression

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.