Back

GTRspmix: Capturing Heterogeneity of Exchangeabilities Across Sites to Improve Protein Phylogenetics

Harada, R.; Susko, E.; Wong, T. K. F.; Banos, H.; Ly-Trong, N.; Lanfear, R.; Theobald, D. L.; Minh, B. Q.; Roger, A. J.

2026-06-18 evolutionary biology
10.64898/2026.06.18.729217 bioRxiv
Show abstract

Site rate and profile mixture models capture the heterogeneity of the amino acid substitution process across sites. However, these models typically use a single matrix of amino acid exchangeabilities and ignore potential heterogeneities of these exchangeabilities across sites. Simply combining multiple exchangeability matrices with rate and profile mixtures leads to a combinatorial explosion of mixture components and a prohibitive increase in free parameters. Here, we introduce GTRspmix, a novel framework that incorporates multiple exchangeability matrices into profile and site rate mixture models while effectively managing model complexity. GTRspmix employs a clustering-based strategy that groups profiles and assigns a distinct exchangeability matrix to each profile cluster. Evaluations using both empirical and simulated datasets demonstrate that GTRspmix fits empirical data significantly better than conventional models, and that overparameterization does not present a problem for sufficiently large alignments. Based on these results, we estimated general-purpose empirical models (SXXpfamCYY series available in IQ-TREE3) from the Pfam database. These general-purpose models not only fit data much better, but they also influence branch length and tree topology estimates, effectively mitigating long-branch attraction artifacts. Because the total number of rate matrices remains manageable, the computational efficiency of the inference is identical to that of conventional profile mixture models (e.g., LG+C60+G4). GTRspmix provides a more realistic and flexible model of protein evolution, offering a robust foundation for the inference of reliable phylogenetic trees.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Molecular Biology and Evolution
542 papers in training set
Top 0.2%
22.2%
2
Systematic Biology
144 papers in training set
Top 0.2%
15.0%
3
Bioinformatics
1204 papers in training set
Top 2%
11.8%
4
PLOS Computational Biology
1863 papers in training set
Top 6%
6.2%
50% of probability mass above
5
Virus Evolution
155 papers in training set
Top 0.5%
4.0%
6
Journal of Computational Biology
48 papers in training set
Top 0.2%
4.0%
7
Protein Science
246 papers in training set
Top 1%
4.0%
8
PLOS ONE
5266 papers in training set
Top 40%
2.8%
9
Genome Biology and Evolution
338 papers in training set
Top 2%
2.4%
10
GENETICS
483 papers in training set
Top 2%
2.1%
11
Nature Communications
5641 papers in training set
Top 42%
2.1%
12
PeerJ
308 papers in training set
Top 5%
1.7%
13
Communications Biology
993 papers in training set
Top 14%
1.7%
14
eLife
5828 papers in training set
Top 52%
1.5%
15
Genome Research
468 papers in training set
Top 4%
1.3%
16
BMC Evolutionary Biology
18 papers in training set
Top 0.2%
1.1%
17
Nature Computational Science
55 papers in training set
Top 1%
1.1%
18
Bioinformatics Advances
203 papers in training set
Top 4%
1.0%
19
Journal of Molecular Evolution
22 papers in training set
Top 0.4%
1.0%
20
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 42%
0.8%
21
Scientific Reports
3612 papers in training set
Top 75%
0.8%
22
Methods in Ecology and Evolution
176 papers in training set
Top 2%
0.6%
23
Journal of Molecular Biology
232 papers in training set
Top 4%
0.6%