Back

Toward Accurate RNA Non-Canonical Structure Prediction: The NC-Bench Benchmark and the NCfold Framework

Zhu, H.; Li, R.; Chang, A.; Li, M.; Chen, H.; Xiong, P.; Zhou, S. K.

2025-11-17 bioinformatics
10.1101/2025.11.16.688746 bioRxiv
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWAccurate prediction of non-canonical (NC) base pairs is a pivotal step toward un-raveling the full functional landscape of RNA. This goal requires both a standardized dataset for evaluation and powerful models that overcome data limitations. Here, we make a dual contribution to this endeavor: the NC-Bench benchmark and the NCfold framework. NC-Bench offers a trusted ground for comparison with 925 sequences and 6,708 annotated NC pairs, defining tasks for edge and orientation prediction. NCfold provides a powerful analytical engine, a closed-loop dual-branch network that synergizes sequence features with structure-aware embeddings from RNA foundation models via our proposed REFweighted self-attention. This iterative design enables robust learning from scarce NC data. On the NC-Bench benchmark, NCfold establishes new state-of-the-art results, significantly outperforming existing methods. Comprehensive analyses validate its design and affirm the critical role of NC-Bench in driving future progress. NC-Bench and NCfold together form a cornerstone for the next generation of RNA structure prediction. The datasets and codes are publicly available at https://github.com/heqin-zhu/NCBench

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.