ICEPIC: A Toolkit to Discover Ice Binding Proteins from Sequence
Zhang, J.; Suresh, S.; Gleizer, S.; Ewens, S.; Venkat, A.; Zulkower, V.; Biernacki, T.; Wen, D.; Li, C.; Eslami, M.; Buckhout-White, S.
Show abstract
Ice binding proteins, such as antifreeze proteins (AFPs) and ice nucleation proteins (INPs), are critical for survival in subzero environments and have wide-ranging applications in biotechnology, agriculture, and materials science. Current discovery methods for these proteins are constrained by low throughput and limited datasets that are not conducive for engineering. Here, we present a high-throughput, sequence-based model that leverages contextual embeddings from protein language models to predict ice binding potential, as well as the expression and activity potential of candidate proteins. Using a curated data corpus of over 18,000 ice binding proteins -- far larger than previous datasets -- we fine-tuned a ProtBERT-based model, achieving 99% accuracy for prediction of ice binding potential. Sensitivity analyses through targeted mutagenesis (alanine and threonine substitutions) confirmed the models biological significance, revealing functionally important residues and sequence patterns. Additionally, we developed an expression prediction model that achieved an R2 score of 0.64 and low false-negative rates in identifying highly expressible candidates in Pichia pastoris. An additional regression model trained to predict ice activity as measured by thermal hysteresis achieved an R2 score of at least 0.79 with a clear difference in prediction between ice binding and non-ice binding proteins. Our toolkit advances the predictive accuracy, interpretability, and scalability of ice binding protein discovery, offering a powerful tool for protein engineering in cold-environment applications. Significance StatementIce binding proteins enable organisms to survive freezing temperatures and are essential for applications in cryopreservation, agriculture, and materials science. However, discovering and engineering these proteins has been limited by small datasets and inadequate predictive tools. We developed a machine learning model trained on over 18,000 ice binding protein sequences to predict not only ice binding potential but also protein activity and expression in engineered hosts. This approach integrates advanced protein language models with biological context, enabling faster, more reliable discovery of ice binding proteins. Our platform advances rational protein design for real-world applications in climate resilience, biotechnology, and beyond.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Species-specific design of artificial promoters by transfer-learning based generative deep-learning model 93%
- iModulonDB: a knowledgebase of microbial transcriptional regulation derived from machine learning 93%
- PanKB: An interactive microbial pangenome knowledgebase for research, biotechnological innovation, and knowledge mining 93%
Similar papers in this journal
- Enzyme structure correlates with variant effect predictability 94%
- A Multi-Layered Computational Structural Genomics Approach Enhances Domain-Specific Interpretation of Kleefstra Syndrome Variants in EHMT1 91%
- DeepNeuropePred: a robust and universal tool to predict cleavage sites from neuropeptide precursors by protein language model 91%
Similar papers in this journal
- Data-Driven Strain Design Using Aggregated Adaptive Laboratory Evolution Mutational Data 92%
- Protein interaction kinetics delimit the performance of phosphorylation-driven protein switches. 92%
- The Synthesis Success Calculator: Predicting the Rapid Synthesis of DNA Fragments with Machine Learning 92%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.