Back

NUMonomer enables accurate and scalable nucleic acid structure prediction from primary sequence alone

si, y.; zhang, s.; chen, l.

2026-07-20 molecular biology
10.64898/2026.07.20.739453 bioRxiv
Show abstract

Accurate and efficient prediction of three-dimensional nucleic acid structures can accelerate functional characterization and enable downstream applications. Recent deep-learning methods have substantially improved nucleic acid structure prediction by incorporating auxiliary inputs such as multiple sequence alignments, secondary-structure annotations, and representations from pretrained language models. However, prediction accuracy remains limited, and generating these auxiliary inputs can be computationally expensive. Here we show that learning the hierarchical organization of experimentally determined structures across multiple scales, from recurring local conformations to global fold topologies, together with exploiting representations shared between RNA and single-stranded DNA, improves model generalization. Guided by these findings, we developed NUMonomer, an end-to-end deep-learning framework trained with input sequences spanning thousands of nucleotides on a joint RNA and single-stranded DNA dataset to predict nucleic acid structures directly from sequence. Despite requiring no auxiliary inputs, NUMonomer matches or outperforms leading prediction methods on benchmarks comprising CASP16 RNA targets and non-redundant sets of experimentally determined RNA and single-stranded DNA structures, with particularly pronounced improvements for longer RNAs. Its efficient and scalable architecture also reduces inference costs by approximately two orders of magnitude relative to the evaluated methods, enabling large-scale structure prediction. Together, these findings provide insight into generalization in biomolecular structure learning and establish NUMonomer as a practical framework for nucleic acid structure prediction.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.