NUMonomer enables accurate and scalable nucleic acid structure prediction from primary sequence alone
si, y.; zhang, s.; chen, l.
Show abstract
Accurate and efficient prediction of three-dimensional nucleic acid structures can accelerate functional characterization and enable downstream applications. Recent deep-learning methods have substantially improved nucleic acid structure prediction by incorporating auxiliary inputs such as multiple sequence alignments, secondary-structure annotations, and representations from pretrained language models. However, prediction accuracy remains limited, and generating these auxiliary inputs can be computationally expensive. Here we show that learning the hierarchical organization of experimentally determined structures across multiple scales, from recurring local conformations to global fold topologies, together with exploiting representations shared between RNA and single-stranded DNA, improves model generalization. Guided by these findings, we developed NUMonomer, an end-to-end deep-learning framework trained with input sequences spanning thousands of nucleotides on a joint RNA and single-stranded DNA dataset to predict nucleic acid structures directly from sequence. Despite requiring no auxiliary inputs, NUMonomer matches or outperforms leading prediction methods on benchmarks comprising CASP16 RNA targets and non-redundant sets of experimentally determined RNA and single-stranded DNA structures, with particularly pronounced improvements for longer RNAs. Its efficient and scalable architecture also reduces inference costs by approximately two orders of magnitude relative to the evaluated methods, enabling large-scale structure prediction. Together, these findings provide insight into generalization in biomolecular structure learning and establish NUMonomer as a practical framework for nucleic acid structure prediction.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- US-align: Universal Structure Alignments of Proteins, Nucleic Acids, and Macromolecular Complexes 98%
- Predicting structures of large protein assemblies using combinatorial assembly algorithm and AlphaFold2 95%
- Direct prediction of intrinsically disordered protein conformational properties from sequence 94%
Similar papers in this journal
Similar papers in this journal
- Predicting RNA 3D structure and conformers using a pre-trained secondary structure model and structure-aware attention 96%
- A 5' UTR Language Model for Decoding Untranslated Regions of mRNA and Function Predictions 96%
- PSICHIC: physicochemical graph neural network for learning protein-ligand interaction fingerprints from sequence data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.