TransfoRNA: Navigating the Uncertainties of Small RNA Annotation with an Adaptive Machine Learning Strategy
Taha, Y.; Jehn, J.; Kahraman, M.; Frank, M.; Heuvelman, M.; Horos, R.; Yau, C.; Steinkraus, B.; Sikosek, T.
Show abstract
Small RNAs hold crucial biological information and have immense diagnostic and therapeutic value. While many established annotation tools focus on microRNAs, there are myriads of other small RNAs that are currently underutilized. These small RNAs can be difficult to annotate, as ground truth is limited and well-established mapping and mismatch rules are lacking. TransfoRNA is a machine learning framework based on Transformers that explores an alternative strategy. It uses common annotation tools to generate a small seed of high-confidence training labels, while then expanding upon those labels iteratively. TransfoRNA learns sequence-specific representations of all RNAs to construct a similarity network which can be interrogated as new RNAs are annotated, allowing to rank RNAs based on their familiarity. While models can be flexibly trained on any RNA dataset, we here present a version trained on TCGA (The Cancer Genome Atlas) small RNA sequences and demonstrate its ability to add annotation confidence to an unrelated dataset, where 21% of previously unannotated RNAs could be annotated. Relative to its training data, TransfoRNA could boost high-confidence annotations in TCGA by [~]50% while providing transparent explanations even for low-confidence ones. It could learn to annotate 97% of isomiRs from just single examples and confidently identify new members of other familiar classes with high accuracy, while reliably rejecting false RNAs. All source code is available at https://github.com/gitHBDX/TransfoRNA and can be executed at Code Ocean (https://codeocean.com/capsule/5415298/). An interactive website is available at www.transforna.com. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=149 SRC="FIGDIR/small/599329v1_ufig1.gif" ALT="Figure 1"> View larger version (53K): org.highwire.dtl.DTLVardef@11bc3d9org.highwire.dtl.DTLVardef@1d6ff01org.highwire.dtl.DTLVardef@1ffac02org.highwire.dtl.DTLVardef@75ccd8_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Benchmarking Pre-trained Genomic Language Models for RNA Sequence-Related Predictive Applications 96%
- Differential Analysis of RNA Structure Probing Experiments at Nucleotide Resolution: Uncovering Regulatory Functions of RNA Structure 96%
- Semi-quantitative detection of pseudouridine modifications and type I/II hypermodifications in human mRNAs using direct and long-read sequencing 96%
Similar papers in this journal
Similar papers in this journal
- Uncalled4 improves nanopore DNA and RNA modification detection via fast and accurate signal alignment 95%
- Single molecule co-occupancy of RNA-binding proteins with an evolved RNA deaminase 95%
- A systematic benchmark of Nanopore long read RNA sequencing for transcript level analysis in human cell lines 94%
Similar papers in this journal
Similar papers in this journal
- Towards In-Silico CLIP-seq: Predicting Protein-RNA Interaction via Sequence-to-Signal Learning 95%
- EvoRMD: Integrating Biological Context and Evolutionary RNA Language Models for Interpretable Prediction of RNA Modifications 95%
- HydraRNA: a hybrid architecture based full-length RNA language model 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.