Back

CoSTAR: Coarse Stem-Topology Alignment of Pseudoknotted RNA Structures by Relation-Constrained Search

Archinuk, F.; Jabbari, H.

2026-05-26 bioinformatics
10.64898/2026.05.22.727263 bioRxiv
Show abstract

RNA structural alignment is a central task in comparative RNA analysis, but many efficient methods achieve tractability by restricting the class of admissible structures, often excluding pseudoknots. This exclusion is limiting for viral and regulatory RNAs, where conserved structure can remain informative even when sequence conservation is weak. We introduce a coarse RNA structural alignment algorithm that aligns secondary structures by searching over partial maps between stems rather than nucleotides. Each input structure is decomposed into stems, annotated with nucleotide-level features, and encoded by pairwise topological relations among stems. Alignment is formulated as a cost-minimizing partial stem map with skip operations, and the search tree is pruned by RNA-specific directionality and topological constraints derived from already aligned stems. For the stated cost function and over the class of injective, direction-preserving, topologically consistent stem maps, the search is exact. This shifts the dominant computational dependence from sequence length to the number and arrangement of stems. We evaluated the method on 2100 pairwise alignments sampled from seven Rfam families spanning 40-224 nucleotides and 2-15 stems. Across these benchmarks, the algorithm returned terminal coarse alignments in which every stem was either matched or skipped. We measured running time and search-tree width to characterize performance on diverse family-to-family comparisons. The experiments also show that ordering the input structures affects efficiency: using the structure with more stems as the search-driving structure reduces tree width. The resulting partial stem map is directly interpretable for RNA annotation and can be projected to nucleotide resolution for downstream sequence-structure analysis. The source code for CoSTAR is available at: https://github.com/TheCOBRALab/CoSTAR

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Nature Communications
5641 papers in training set
Top 11%
16.6%
2
Bioinformatics
1204 papers in training set
Top 2%
12.6%
3
Nucleic Acids Research
1281 papers in training set
Top 2%
9.4%
4
Nature Methods
385 papers in training set
Top 1%
7.7%
5
Nature Biotechnology
172 papers in training set
Top 0.7%
5.4%
50% of probability mass above
6
Genome Biology
637 papers in training set
Top 2%
5.4%
7
RNA
189 papers in training set
Top 0.4%
4.7%
8
PLOS Computational Biology
1863 papers in training set
Top 9%
3.9%
9
Cell Systems
201 papers in training set
Top 1%
3.9%
10
NAR Genomics and Bioinformatics
242 papers in training set
Top 1%
3.2%
11
Bioinformatics Advances
203 papers in training set
Top 2%
2.7%
12
Algorithms for Molecular Biology
17 papers in training set
Top 0.1%
2.6%
13
Genome Research
468 papers in training set
Top 3%
2.4%
14
Nature Genetics
286 papers in training set
Top 3%
1.6%
15
BMC Bioinformatics
457 papers in training set
Top 5%
1.1%
16
eLife
5828 papers in training set
Top 59%
1.1%
17
Scientific Reports
3612 papers in training set
Top 69%
1.0%
18
Molecular Biology and Evolution
542 papers in training set
Top 4%
1.0%
19
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 37%
1.0%
20
Nature
645 papers in training set
Top 10%
0.9%
21
Biophysical Journal
631 papers in training set
Top 5%
0.8%
22
Nature Machine Intelligence
70 papers in training set
Top 3%
0.6%