Prediction of Protein Half-lives from Amino Acid Sequences by Protein Language Models
Sagawa, T.; Kanao, E.; Ogata, K.; Imami, K.; Ishihama, Y.
Show abstract
We developed a protein half-life prediction model, PLTNUM, based on a protein language model using an extensive dataset of protein sequences and protein half-lives from the NIH3T3 mouse embryo fibroblast cell line as a training set. PLTNUM achieved an accuracy of 71% on validation data and showed robust performance with an ROC of 0.73 when applied to a human cell line dataset. By incorporating Shapley Additive Explanations (SHAP) into PLTNUM, we identified key factors contributing to shorter protein half-lives, such as cysteine-containing domains and intrinsically disordered regions. Using SHAP values, PLTNUM can also predict potential degron sequences that shorten protein half-lives. This model provides a platform for elucidating the sequence dependency of protein half-lives, while the uncertainty in predictions underscores the importance of biological context in influencing protein half-lives.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Generalizable and scalable protein stability prediction with rewired protein generative models 96%
- High-accuracy protein complex structure modeling based on sequence-derived structure complementarity 96%
- Retention Time Prediction Using Neural Networks Increases Identifications in Crosslinking Mass Spectrometry 96%
Similar papers in this journal
- Direct prediction of intrinsically disordered protein conformational properties from sequence 96%
- Sliding Window INteraction Grammar (SWING): a generalized interaction language model for peptide and protein interactions 95%
- Predicting structures of large protein assemblies using combinatorial assembly algorithm and AlphaFold2 94%
Similar papers in this journal
Similar papers in this journal
- The amino acid sequence determines protein abundance through its conformational stability and reduced synthesis cost. 97%
- Learning deep representations of enzyme thermal adaptation 96%
- COLLAPSE: A representation learning framework for identification and characterization of protein structural sites 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.