OPUS-Design: Designing Protein Sequence from Backbone Structure with 3DCNN and Protein Language Model
Xu, G.; Yang, Y.; Zhang, Y.; Wang, Q.; Ma, J.
Show abstract
Protein sequence design, also known as protein inverse folding, is a crucial task in protein engineering and design. Despite the recent advancements in this field, which have facilitated the identification of amino acid sequences based on backbone structures, achieving higher levels of accuracy in sequence recovery rates remains challenging. It this study, we introduce a two-stage protein sequence design method named OPUS-Design. Our evaluation on recently released targets from CAMEO and CASP15 shows that OPUS-Design significantly surpasses several other leading methods on both monomer and oligomer targets in terms of sequence recovery rate. Furthermore, by utilizing its finetune version OPUS-Design-ft and our previous work OPUS-Mut, we have successfully designed a thermal-tolerant double-point mutant of T4 lysozyme that demonstrates a residual enzyme activity exceeding that of the wild-type T4 by more than twofold when both are subjected to extreme heat treatment at 70{degrees}C. Importantly, this accomplishment is achieved through the experimental verification of less than 10 mutant candidates, thus significantly alleviating the burden of experimental verification process.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- CONSTRUCT: an algorithmic tool for identifying functional or structurally important regions in protein tertiary structure 96%
- VoroCNN: Deep convolutional neural network built on 3D Voronoi tessellation of protein structures 96%
- TemStaPro: protein thermostability prediction using sequence representations from protein language models 95%
Similar papers in this journal
- PremPS: Predicting the Effects of Single Mutations on Protein Stability 96%
- Predicting changes in protein thermodynamic stability upon point mutation with deep 3D convolutional neural networks 96%
- Deducing high-accuracy protein contact-maps from a triplet of coevolutionary matrices through deep residual convolutional networks 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.