TISCalling: Leveraging Machine Learning to Identify Translational Initiation Sites in Plants and Viruses
Yen, M.-R.; Cheng, C.-Y.; Wu, T.-Y.; Liu, M.-J.
Show abstract
The recognition of translational initiation sites (TISs) offers complementary insights into identifying genes encoding novel proteins or small peptides. Conventional computational methods primarily identify Ribo-seq-supported TISs and lack the capacity of systematical and global identification of TIS, especially for non-AUG sites in plants. Additionally, these methods are often unsuitable for evaluating the importance of mRNA sequence features for TIS determination. In this study, we present TISCalling, a robust framework that combines machine learning (ML) models and statistical analysis to identify and rank novel TISs across eukaryotes. TISCalling generalized and ranks important features common to multiple plant and mammalian species while identifying kingdom-specific features such as mRNA secondary structures and G-contents. Furthermore, TISCalling achieved high predictive power for identifying novel viral TISs. Importantly, TISCalling provides prediction scores for putative TIS along plant transcripts, enabling prioritization of those of interest for further validation. We offer TISCalling as a command-line-based package [https://github.com/yenmr/TIScalling], capable of generating prediction models and identifying key sequence features. Additionally, we provide web tools [https://predict.southerngenomics.org/TIScalling] for visualizing pre-computed potential TISs, making it accessible to users without programming experience. The TISCalling framework offers a sequence-aware and interpretable approach for decoding genome sequences and exploring functional proteins in plants and viruses.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Full-length isoform constructor (FLIC) - a tool for isoform discovery based on long reads 93%
- RiboFlow, RiboR and RiboPy: An ecosystem for analyzing ribosome profiling data at read length resolution 93%
- HiTea: a computational pipeline to identify non-reference transposable element insertions in Hi-C data 93%
Similar papers in this journal
- DiffSegR: An RNA-Seq data driven method for differential expression analysis using changepoint detection 92%
- Accurate prediction of cis-regulatory modules reveals a prevalent regulatory genome of humans 92%
- GeneMark-EP and -EP+: eukaryotic gene prediction with self-training in the space of genes and proteins 92%
Similar papers in this journal
- At-RS31 orchestrates hierarchical cross-regulation of splicing factors and integrates alternative splicing with TOR-ABA pathways 93%
- A high-quality genome of the mangrove Aegiceras corniculatum aids investigation of molecular adaptation to intertidal environments 92%
- Identification of new marker genes from plant single-cell RNA-seq data using interpretable machine learning methods 92%
Similar papers in this journal
- Genes with 5' terminal oligopyrimidine tracts preferentially escape global suppression of translation by the SARS-CoV-2 NSP1 protein 93%
- IsoformMapper: A Web Application for Protein-Level Comparison of Splice Variants through Structural Community Analysis 93%
- Extensible benchmarking of methods that identify and quantify polyadenylation sites from RNA-seq data 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.