Back

Quantifying and Predicting the Difficulty of Multiple Sequence Alignment with AlDiScore

Bodynek, M.; Martin-Fernandez, L.; Bettisworth, B.; Haag, J.; Stamatakis, A.

2026-06-02 bioinformatics
10.64898/2026.05.29.727837 bioRxiv
Show abstract

Multiple Sequence Alignment (MSA) constitutes an important and frequent operation in molecular sequence data analysis. There exist numerous tools, algorithms, and criteria to infer an MSA. This plethora of available approaches to MSA may induced an ensemble of divergent MSAs for the same underlying unaligned sequence set. Even a single MSA tool may infer distinct MSAs when varying the input parameters. Hence, when using a diversified set of MSA algorithms and parameterizations, the observed dispersion within an MSA ensemble expresses the difficulty of inferring a robust alignment. We refer to this notion as MSA difficulty. As downstream analyses heavily rely on the MSA, characterizing MSA difficulty for a given unaligned sequence set is critical. Initially, we show that measures of dispersion within diversified MSA ensembles can reliably predict MSA difficulty. We then assess the adequacy of these measures by computing the average reference-based distance between the MSAs in the MSA ensemble and its corresponding structural reference MSA and subsequently comparing this distance to the corresponding reference-free average distance over all MSA pairs in the ensemble. We find that Blackburne and Whelans dpos alignment metric is most appropriate as its reference-free [Formula] counterpart most accurately approximates the reference-based difficulty computed on BAliBASE reference data. We therefore use [Formula] to quantify MSA difficulty on a scale from 0 (easy) to 1 (difficult). Next, we introduce the AlDiScore open-source tool, which uses machine learning to directly and reliably predict reference-free difficulty scores from unaligned sequence sets to completely omit expensive MSA computations. The underlying regression model relies upon a large set of features, including sampling-based measures of transitive consistency. We trained our AlDiScore model on a diverse collection of empirical datasets from BAliBASE, TreeBASE, and published studies. Subsequently, we demonstrate that AlDiScore attains an R2 of 0.89 and of 0.84 on unseen AA and DNA sequence sets extracted from the PANDIT v17 database. Finally, we show that there is no correlation between MSA difficulty and the corresponding phylogenetic difficulty of the respective MSA.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 1%
16.7%
2
Molecular Biology and Evolution
542 papers in training set
Top 0.9%
7.7%
3
PLOS Computational Biology
1863 papers in training set
Top 5%
7.7%
4
Journal of Chemical Information and Modeling
238 papers in training set
Top 0.9%
5.4%
5
Systematic Biology
144 papers in training set
Top 0.3%
4.8%
6
Proteins: Structure, Function, and Bioinformatics
88 papers in training set
Top 0.2%
4.8%
7
BMC Bioinformatics
457 papers in training set
Top 2%
4.3%
50% of probability mass above
8
Biophysical Journal
631 papers in training set
Top 2%
3.4%
9
NAR Genomics and Bioinformatics
242 papers in training set
Top 1%
3.2%
10
Bioinformatics Advances
203 papers in training set
Top 2%
3.1%
11
Briefings in Bioinformatics
354 papers in training set
Top 3%
2.6%
12
PLOS ONE
5266 papers in training set
Top 43%
2.4%
13
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 23%
2.3%
14
Algorithms for Molecular Biology
17 papers in training set
Top 0.1%
1.9%
15
Scientific Reports
3612 papers in training set
Top 55%
1.7%
16
Journal of Computational Biology
48 papers in training set
Top 0.6%
1.7%
17
Protein Science
246 papers in training set
Top 2%
1.5%
18
Nature Communications
5641 papers in training set
Top 48%
1.5%
19
eLife
5828 papers in training set
Top 53%
1.4%
20
Nucleic Acids Research
1281 papers in training set
Top 11%
1.1%
21
PeerJ
308 papers in training set
Top 8%
1.1%
22
Frontiers in Bioinformatics
49 papers in training set
Top 0.9%
1.1%
23
Nature Computational Science
55 papers in training set
Top 1%
1.1%
24
Nature Methods
385 papers in training set
Top 6%
1.0%
25
Journal of Chemical Theory and Computation
140 papers in training set
Top 1%
1.0%
26
Journal of Computational Chemistry
13 papers in training set
Top 0.3%
0.8%
27
Journal of Molecular Biology
232 papers in training set
Top 4%
0.8%
28
Cell Systems
201 papers in training set
Top 5%
0.8%
29
GENETICS
483 papers in training set
Top 4%
0.8%
30
ACS Omega
105 papers in training set
Top 4%
0.8%