Back

Spatial-neighbour encoding enables fast RNA 3D structure search

Wang, D.; Jin, J.; Qiao, J.; Wei, L.; Wu, S.; Liu, Q.

2026-04-22 bioinformatics
10.64898/2026.04.19.719441 bioRxiv
Show abstract

Experimental and predicted RNA three-dimensional structures are expanding rapidly, but RNA structure search still lacks a compact residue-level representation that supports database-scale comparison. Using family-held-out ablations across the currently available experimental RNA structure collection, we found that spatial-neighbour features are markedly more informative for family-level discrimination than conventional backbone and base descriptors. Building on this result, we developed RiboSeek, a search framework based on a 20-letter geometric alphabet (RS-20), an 80-letter structure-and-base composite alphabet (RS-80). Across family-level classification and retrieval benchmarks, RS-80 delivered the strongest overall performance, whereas RS-20 most closely tracked US-align TM-score, indicating better preservation of geometric similarity. RiboSeek searches the full experimental RNA structure database in 204 ms per query and can be applied to predicted RNA structure libraries to prioritize candidate structural relationships for downstream analysis.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Nucleic Acids Research
1128 papers in training set
Top 0.1%
31.6%
2
Nature Communications
4913 papers in training set
Top 24%
8.1%
3
PLOS Computational Biology
1633 papers in training set
Top 6%
6.0%
4
Cell Systems
167 papers in training set
Top 2%
6.0%
50% of probability mass above
5
Nature Methods
336 papers in training set
Top 3%
3.8%
6
Nature Biotechnology
147 papers in training set
Top 2%
3.8%
7
Bioinformatics Advances
184 papers in training set
Top 2%
3.4%
8
Proteins: Structure, Function, and Bioinformatics
82 papers in training set
Top 0.2%
3.1%
9
PLOS ONE
4510 papers in training set
Top 42%
2.9%
10
Bioinformatics
1061 papers in training set
Top 6%
2.9%
11
NAR Genomics and Bioinformatics
214 papers in training set
Top 1%
2.6%
12
Briefings in Bioinformatics
326 papers in training set
Top 3%
2.3%
13
Genome Research
409 papers in training set
Top 2%
1.7%
14
Journal of Chemical Information and Modeling
207 papers in training set
Top 2%
1.6%
15
Communications Biology
886 papers in training set
Top 10%
1.6%
16
Structure
175 papers in training set
Top 2%
1.2%
17
Scientific Reports
3102 papers in training set
Top 70%
0.9%
18
Journal of Molecular Biology
217 papers in training set
Top 3%
0.9%
19
RNA
169 papers in training set
Top 0.4%
0.9%
20
Genomics, Proteomics & Bioinformatics
171 papers in training set
Top 5%
0.9%
21
Genome Biology
555 papers in training set
Top 7%
0.8%
22
Genome Medicine
154 papers in training set
Top 9%
0.7%
23
BMC Bioinformatics
383 papers in training set
Top 7%
0.7%
24
Computational and Structural Biotechnology Journal
216 papers in training set
Top 10%
0.7%
25
Cell Genomics
162 papers in training set
Top 8%
0.6%