Large datasets and machine learning models fail to capture extremophile enzyme melting and optimum temperatures
Gault, S.
Show abstract
Organisms and their enzymes adapt to environmental temperatures, such that thermophilic enzymes exhibit high melting and optimum temperatures while psychrophilic enzymes exhibit low values for both. It has been proposed that the gap between an enzymes optimum temperature and its melting temperature, the temperature gap, is characteristically large in psychrophiles, implying that the loss of activity above the optimum is decoupled from global protein stability. The evidence for this relies on a small number of characterised enzymes, leaving the prevalence of large temperature gaps amongst psychrophiles unknown. We asked whether the machine-learning predictors and large datasets now available could test this at scale. We find that they cannot: predictors of melting and optimum temperature fail systematically at the thermal extremes, assigning the majority of thermophilic enzymes with optimum temperatures that exceed their melting temperatures, which is biophysically implausible, and consistently underpredict the stability of (hyper)thermophiles. This stems from training data that is both error-laden, as we demonstrate for widely used optimum-temperature records, and overwhelmingly biased toward mesophiles, which regresses predictions for cold and heat adapted enzymes toward mesophilic values. Consequently, current computational tools cannot establish how prevalent the psychrophilic temperature gap is. We argue that proteome-scale measurement of extremophile enzyme thermal behaviour, integrated as curated training data, is required to determine whether trends from small studies extend across the diversity of life.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Computation of condition-dependent proteome allocation reveals variability in the macro and micro nutrient requirements for growth 93%
- Towards a comprehensive view of the pocketome universe - biological implications and algorithmic challenges. 93%
- Protein Stability Prediction by Fine-tuning a Protein Language Model on a Mega-scale Dataset 93%
Similar papers in this journal
- Computational investigation of missense somatic mutations in cancer and potential links to pH-dependence and proteostasis 91%
- dagLogo: an R/Bioconductor package for identifying and visualizing differential amino acid group usage in proteomics data 91%
- Phylogenetic profiling in eukaryotes: The effect of species, orthologous group, and interactome selection on protein interaction prediction 91%
Similar papers in this journal
- Structure and function of aerotolerant, multiple-turnover THI4 thiazole synthases 90%
- Itaconate utilisation by the human pathogen Pseudomonas aeruginosa requires uptake via the IctPQM TRAP transporter 90%
- Flexibility of Short-chain dehydrogenase is interconnected to its promiscuity for the reduction of multiple ketone intermediates 89%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.