Back

Homology-aware cross-validation strategies for generalization assessment in RNA structure prediction

Bugnon, L.; Kulemeyer, G.; Gerard, M.; Di Persia, L.; Stegmayer, G.; Milone, D. H.

2026-06-29 bioinformatics
10.64898/2026.06.28.735057 bioRxiv
Show abstract

RNA secondary structure prediction is a fundamental challenge in bioinformatics, essential for understanding the functional roles of non-coding RNAs. Recently, deep learning models have transformed the field with impressive results, leading to critical discussions regarding the validity of current cross-validation strategies. On the one hand, traditional random partitioning yields overop-timistic results due to data leakage from uncontrolled homology. On the other hand, removing from the training set all sequences that exhibit even the slightest resemblance to the testing sequences penalizes learning-based methods by requiring generalization to completely out-of-distribution sequences. While it is very simple to remove sequences and retrain a machine learned model, it is very difficult to remove the experimental data used for parameter tuning and the sequences used for the development of classical thermodynamic methods. Thus, these methods often benefit from an implicit knowledge leakage. In this work we critically review existing cross-validation strategies for RNA secondary structure prediction: random splitting, clustering-based splitting, and leaving one RNA family out for testing. We analyze the advantages and limitations of each strategy, also expanding them towards the future directions to ensure fair comparisons across the full range of sequence similarities, with the same rigor for both classical and learning-based methods.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
Nucleic Acids Research
1281 papers in training set
Top 2%
11.7%
2
PLOS Computational Biology
1863 papers in training set
Top 3%
10.4%
3
Bioinformatics
1204 papers in training set
Top 3%
9.6%
4
Journal of Chemical Theory and Computation
140 papers in training set
Top 0.2%
9.5%
5
Briefings in Bioinformatics
354 papers in training set
Top 1%
7.1%
6
Computational and Structural Biotechnology Journal
242 papers in training set
Top 0.4%
6.1%
50% of probability mass above
7
RNA
189 papers in training set
Top 0.5%
4.2%
8
NAR Genomics and Bioinformatics
242 papers in training set
Top 1%
4.2%
9
Journal of Chemical Information and Modeling
238 papers in training set
Top 1%
2.6%
10
Bioinformatics Advances
203 papers in training set
Top 2%
2.3%
11
Journal of Computational Chemistry
13 papers in training set
Top 0.1%
1.7%
12
BMC Bioinformatics
457 papers in training set
Top 4%
1.7%
13
Genomics, Proteomics & Bioinformatics
16 papers in training set
Top 0.1%
1.7%
14
Scientific Reports
3612 papers in training set
Top 57%
1.6%
15
Nature Communications
5641 papers in training set
Top 47%
1.6%
16
Proteins: Structure, Function, and Bioinformatics
88 papers in training set
Top 0.9%
1.4%
17
International Journal of Molecular Sciences
494 papers in training set
Top 12%
1.1%
18
Frontiers in Bioinformatics
49 papers in training set
Top 0.9%
1.1%
19
Biophysical Journal
631 papers in training set
Top 4%
1.1%
20
Molecular Biology and Evolution
542 papers in training set
Top 4%
1.1%
21
Journal of Computational Biology
48 papers in training set
Top 0.9%
1.0%
22
Genome Biology
637 papers in training set
Top 8%
0.9%
23
PLOS ONE
5266 papers in training set
Top 62%
0.8%
24
Genes
144 papers in training set
Top 4%
0.8%
25
RNA Biology
78 papers in training set
Top 1%
0.8%
26
BMC Genomics
406 papers in training set
Top 10%
0.6%
27
Cell Reports Methods
165 papers in training set
Top 5%
0.6%
28
Computational Biology and Chemistry
28 papers in training set
Top 1%
0.6%
29
ACS Omega
105 papers in training set
Top 4%
0.6%