Back

PKProbDesign: RNA inverse folding including pseudoknots by optimizing thermodynamic folding probability

Otagaki, T.; Iwakiri, J.; terai, g.; Asai, K.; Sato, K.

2026-07-11 bioinformatics
10.64898/2026.07.09.736945 bioRxiv
Show abstract

MotivationRNA inverse folding, the design of RNA sequences that fold into specified target structures, is a central problem in RNA design, with applications in functional RNA engineering, synthetic biology, and nucleic-acid therapeutics. This task becomes especially challenging for pseudoknotted target structures because pseudoknots disrupt the nested structure assumed by standard thermodynamic folding models. Existing pseudoknot inverse-folding methods often rely on structure-predictor-based objectives. Direct optimization of the thermodynamic folding probability of a specified pseudoknotted target remains limited. This requires an evaluator that can assign target-specific folding probabilities within a pseudoknot-aware ensemble and can be used as an optimization signal. ResultsWe present PKProbDesign, a sampling-based inverse-folding framework that directly optimizes a thermodynamic folding-probability objective for pseudoknotted targets. For each target, candidate sequences are scored by combining the folding probability of a pseudoknot-free scaffold with the conditional folding probability of the remaining extension component. On 354 PseudoBase++ targets, PKProbDesign achieved the highest folding probability on 221 targets, compared with 117 for DesiRNA and 16 for MODENA. ConclusionsPKProbDesign demonstrates that pseudoknot inverse folding can be formulated around target folding probabilities rather than structure-prediction agreement alone. By combining scaffold decomposition with HFold/CParty-consistent conditional-ensemble evaluation, the method provides a practical probability-based framework for designing sequences for density-2 pseudoknotted targets. AvailabilityThe source code of PKProbDesign is available at https://github.com/TakumiOtagaki/PKProbDesign.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 0.2%
45.0%
2
Bioinformatics Advances
203 papers in training set
Top 0.4%
7.8%
50% of probability mass above
3
Nucleic Acids Research
1281 papers in training set
Top 4%
4.8%
4
PLOS Computational Biology
1863 papers in training set
Top 9%
4.0%
5
NAR Genomics and Bioinformatics
242 papers in training set
Top 1%
4.0%
6
Briefings in Bioinformatics
354 papers in training set
Top 2%
3.5%
7
BMC Bioinformatics
457 papers in training set
Top 3%
2.7%
8
Journal of Chemical Theory and Computation
140 papers in training set
Top 0.6%
2.4%
9
RNA
189 papers in training set
Top 0.8%
2.1%
10
Molecular Therapy Nucleic Acids
39 papers in training set
Top 0.4%
2.1%
11
Journal of Chemical Information and Modeling
238 papers in training set
Top 2%
2.1%
12
Genome Biology
637 papers in training set
Top 6%
1.7%
13
Nature Communications
5641 papers in training set
Top 46%
1.7%
14
Computational and Structural Biotechnology Journal
242 papers in training set
Top 5%
1.1%
15
RNA Biology
78 papers in training set
Top 0.9%
1.1%
16
Cell Reports Methods
165 papers in training set
Top 3%
1.0%
17
Scientific Reports
3612 papers in training set
Top 70%
1.0%
18
BMC Genomics
406 papers in training set
Top 8%
0.8%
19
The Journal of Physical Chemistry B
167 papers in training set
Top 2%
0.8%
20
Journal of Computational Chemistry
13 papers in training set
Top 0.3%
0.6%
21
Journal of Molecular Biology
232 papers in training set
Top 4%
0.6%
22
PLOS ONE
5266 papers in training set
Top 65%
0.6%