APHIX: Analysis Pipeline for HIV-1 Isoform eXploration Using Long-read RNA Sequencing Data
Albert, J. L.; Gallardo, C. M.; Torbett, B. E.
Show abstract
HIV-1 uses 4 major splice donors and 8 major splice acceptors as well as dozens of minor, cryptic, and uncharacterized splice sites to produce over one hundred distinct transcript isoforms from a single 9.2 kb genome. As a result, existing bioinformatic pipelines struggle to accurately analyze spliced HIV sequences due to the complex nature of HIV alternative splicing compared to human mRNA splicing. Previous approaches to identify HIV isoforms from long-read sequencing data used pipelines that are not publicly available, are convoluted to operate, or are locked into a specific HIV strain, which limits their wide adoption to other experimental designs or systems. To address this gap, we have developed a bioinformatic pipeline called APHIX that fully automates spliced isoform assignment, splice site usage quantification, and non-coding exon detection. APHIX takes a FASTQ/A of long-read transcripts and a HIV genome reference sequence and fully automates HIV isoform analysis. APHIX calculates splice site usage counts and percentages for each donor and acceptor site and their pairwise combinations, accurately assigns isoforms, and automatically identifies transcripts containing non-coding exons. APHIX is compatible with long-reads sequences generated from multiple platforms and library preps, including direct DNA and RNA sequencing. APHIX can also be adapted to multiple HIV-1 clades and strains by providing the appropriate reference sequence during bioinformatic processing. Overall, APHIX enables comprehensive processing of spliced sequences with reproducible results in a manner that is faster and easier to run compared to other methods.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Selective Ablation of 3' RNA ends and Processive RTs Facilitate Direct cDNA Sequencing of Full-length Host Cell and Viral Transcripts 94%
- Co-variation of viral recombination with single nucleotide variants during virus evolution revealed by CoVaMa 93%
- Nanopore ReCappable Sequencing maps SARS-CoV-2 5' capping sites and provides new insights into the structure of sgRNAs 92%
Similar papers in this journal
- A Comprehensive Annotation of Conserved Protein Domains in Human Endogenous Retroviruses 92%
- Kmerator Suite: design of specific k-mer signatures andautomatic metadata discovery in large RNA-Seq datasets. 92%
- High-resolution HIV-1 m6A epitranscriptome reveals isoform-dependent methylation clusters and unique 2-LTR transcript modifications 92%
Similar papers in this journal
Similar papers in this journal
- VIRUSBreakend: Viral Integration Recognition Using Single Breakends 95%
- V-pipe: a computational pipeline for assessing viral genetic diversity from high-throughput sequencing data 92%
- InterARTIC: an interactive web application for whole-genome nanopore sequencing analysis of SARS-CoV-2 and other viruses 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.