Circular RNA identification using a genomic language model and a small number of authenticated examples
Li, K.; Wang, W.; Jiang, J.; Deng, J.; Zhang, J.; Qiu, S.; Zhang, W.
Show abstract
Genomic language models (gLMs) hold great promise for deciphering biological sequences, yet their effectiveness is hindered by the limited number of experimentally verified examples available for model training, a ubiquitous bottleneck for supervised machine learning. To overcome this challenge, we developed circFormer, the first gLM-driven approach for circular RNA (circRNA) identification. circFormer integrates curriculum learning with gLM fine-tuning: a Nucleotide Transformer model is first trained on a small set of validated circRNAs, the resulting model is used as a teacher to score [~]2.3 million noisy candidates, and the model is then fine-tuned with the noisy candidates along with their scores to improve prediction. Operating either as a standalone predictor or as a filter for existing pipelines, circFormer consistently outperformed traditional machine-learning approaches in accuracy and robustness. Among 50 circFormer-selected candidates that were overlooked by most existing tools, experimental validation using RNase R digestion and RT-qPCR confirmed 94.1% (32/34) of the evaluable candidates as genuine circRNAs. To enhance interpretability, we introduced a model-agnostic, dual-level explainable AI strategy that reveals mechanistic signatures of circRNA formation. circFormer provides a scalable, interpretable, and generalizable framework for converting noisy high-throughput data into reliable functional annotations, highlighting a practical path forward for gLM-based genomics in data-scarce settings.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- CapTrap-Seq: A platform-agnostic and quantitative approach for high-fidelity full-length RNA transcript sequencing 95%
- Benchmarking Pre-trained Genomic Language Models for RNA Sequence-Related Predictive Applications 95%
- Demultiplexing and barcode-specific adaptive sampling for nanopore direct RNA sequencing 95%
Similar papers in this journal
Similar papers in this journal
- Uncalled4 improves nanopore DNA and RNA modification detection via fast and accurate signal alignment 95%
- The Nucleotide Transformer: Building and Evaluating Robust Foundation Models for Human Genomics 95%
- SQANTI3: curation of long-read transcriptomes for accurate identification of known and novel isoforms 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.