PLM_Sol: predicting protein solubility by benchmarking multiple protein language models with the updated Escherichia coli protein solubility dataset
Zhang, X.; Hu, X.; Zhang, T.; Ling, Y.; Liu, C.; Xu, N.; Wang, H.; Sun, W.
Show abstract
Protein solubility plays a crucial role in various biotechnological, industrial and biomedical applications. With the reduction in sequencing and gene synthesis costs, the adoption of high-throughput experimental screening coupled with tailored bioinformatic prediction has witnessed a rapidly growing trend for the development of novel functional enzymes of interest (EOI). High protein solubility rates are essential in this process and accurate prediction of solubility is a challenging task. As deep learning technology continues to evolve, attention-based protein language models (PLMs) can extract intrinsic information from protein sequences to a greater extent. Leveraging these models along with the increasing availability of protein solubility data inferred from structural database like the Protein Data Bank (PDB), holds great potential to enhance the prediction of protein solubility. In this study, we curated an Updated Escherichia coli (E.coli) protein Solubility DataSet (UESolDS) and employed a combination of multiple PLMs and classification layers to predict protein solubility. The resulting best-performing model, named Protein Language Model-based protein Solubility prediction model (PLM_Sol), demonstrated significant improvements over previous reported models, achieving a notable 5.7% increase in accuracy, 9% increase in F1_score, and 10.4% increase in MCC score on the independent test set. Moreover, additional evaluation utilizing our in-house synthesized protein resource as test data, encompassing diverse types of enzymes, also showcased the superior performance of PLM_Sol. Overall, PLM_Sol exhibited consistent and promising performance across both independent test set and experimental set, thereby making it well-suited for facilitating large-scale EOI studies. PLM_Sol is available as a standalone program and as an easy-to-use model at https://zenodo.org/doi/10.5281/zenodo.10675340.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Enzyme structure correlates with variant effect predictability 95%
- A Multi-Layered Computational Structural Genomics Approach Enhances Domain-Specific Interpretation of Kleefstra Syndrome Variants in EHMT1 93%
- Electron microscopy reveals toroidal shape of master neuronal cell differentiator REST - RE1-Silencing Transcription factor 92%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- CONSTRUCT: an algorithmic tool for identifying functional or structurally important regions in protein tertiary structure 95%
- TemStaPro: protein thermostability prediction using sequence representations from protein language models 95%
- CATHe: Detection of remote homologues for CATH superfamilies using embeddings from protein language models 94%
Similar papers in this journal
- SenseNet, a tool for analysis of protein structure networks obtained from molecular dynamics simulations 93%
- Expression, purification and preliminary pharmacological characterization of the Plasmodium falciparum membrane-bound pyrophosphatase type 1 93%
- Biophysical and biochemical evidence for the role of acetate kinases (AckAs) in an acetogenic pathway in pathogenic spirochetes 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.