Cross-species vs species-specific models for protein melting temperature prediction
Garcia Lopez, S.; Salomon, J.; Boomsma, W.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWProtein melting temperatures are important proxies for stability, and frequently probed in protein engineering campaigns, for instance for enzyme discovery and protein optimization. With the emergence of large datasets of melting temperatures for diverse natural proteins, it has become possible to train models to predict this quantity, and the literature has reported impressive performance values in terms of Spearman rho. The high correlation scores suggest that it should be possible to accurately predict melting temperature changes in engineered variants, and to reliably identify naturally thermostable proteins. However, in practice, results in these settings are often disappointing. In this paper, we explore this apparent discrepancy. We show that Spearman rho over cross-species data gives an overly optimistic impression of prediction performance, and that this metric reflects the ability to distinguish global differences in amino acid composition between species, rather than the specific effects of genetic variation. We proceed by investigating whether cross-species training on melting temperature is beneficial at all, compared to training specific models for each species. We address this question using four different transfer-learning approaches and a fine-tuning procedure. Surprisingly, we consistently find no benefit of cross-species training. We conclude that 1) current models for supervised prediction of melting temperature perform substantially worse than the literature suggests, and 2) that reliable transfer across species is still a challenging problem. An implementation of this work is available at https://github.com/deltadedirac/thermocontrast_tm
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- MULAN: Multimodal Protein Language Model for Sequence and Structure Encoding 95%
- DeepRank-GNN-esm: A Graph Neural Network for Scoring Protein-Protein Models using Protein Language Model 95%
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.