PolyAMiner-Bulk: A Machine Learning Based Bioinformatics Algorithm to Infer and Decode Alternative Polyadenylation Dynamics from bulk RNA-seq data
Jonnakuti, V. S.; Wagner, E. J.; Maletic-Savatic, M.; Liu, Z.; Yalamanchili, H. K.
Show abstract
More than half of human genes exercise alternative polyadenylation (APA) and generate mRNA transcripts with varying 3 untranslated regions (UTR). However, current computational approaches for identifying cleavage and polyadenylation sites (C/PASs) and quantifying 3UTR length changes from bulk RNA-seq data fail to unravel tissue- and disease-specific APA dynamics. Here, we developed a next-generation bioinformatics algorithm and application, PolyAMiner-Bulk, that utilizes an attention-based machine learning architecture and an improved vector projection-based engine to infer differential APA dynamics accurately. When applied to earlier studies, PolyAMiner-Bulk accurately identified more than twice the number of APA changes in an RBM17 knockdown bulk RNA-seq dataset compared to current generation tools. Moreover, on a separate dataset, PolyAMiner-Bulk revealed novel APA dynamics and pathways in scleroderma pathology and identified differential APA in a gene that was identified as being involved in scleroderma pathogenesis in an independent study. Lastly, we used PolyAMiner-Bulk to analyze the RNA-seq data of post-mortem prefrontal cortexes from the ROSMAP data consortium and unraveled novel APA dynamics in Alzheimers Disease. Our method, PolyAMiner-Bulk, creates a paradigm for future alternative polyadenylation analysis from bulk RNA-seq data.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Shiba: A versatile computational method for systematic identification of differential RNA splicing across platforms 97%
- Quality-controlled R-loop meta-analysis reveals the characteristics of R-Loop consensus regions 96%
- TREND-DB - A Transcriptome-wide Atlas of the Dynamic Landscape of Alternative Polyadenylation 96%
Similar papers in this journal
- The long and the short of it: unlocking nanopore long-read RNA sequencing data with short-read tools 96%
- A computational method for direct imputation of cell type-specific expression profiles and cellular compositions from bulk-tissue RNA-Seq in brain disorders 95%
- Specifying cellular context of transcription factor regulons for exploring context-specific gene regulation programs 95%
Similar papers in this journal
- Leveraging omic features with F3UTER enables identification of unannotated 3'UTRs for synaptic genes 95%
- Semi-quantitative detection of pseudouridine modifications and type I/II hypermodifications in human mRNAs using direct and long-read sequencing 95%
- G4mer: An RNA language model for transcriptome-wide identification of G-quadruplexes and disease variants from population-scale genetic data 95%
Similar papers in this journal
- Predicting unrecognized enhancer-mediated genome topology by an ensemble machine learning model 95%
- Biosurfer for systematic tracking of regulatory mechanisms leading to protein isoform diversity 95%
- Classification and clustering of RNA crosslink-ligation data reveal complex structures and homodimers 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.