Estimating the Designability of Protein Structures
Pan, F.; Zhang, Y.; Liu, X.; Zhang, J.
Show abstract
The total number of amino acid sequences that can fold to a target protein structure, known as "designability", is a fundamental property of proteins that contributes to their structure and function robustness. The highly designable structures always have higher thermodynamic stability, mutational stability, fast folding, regular secondary structures, and tertiary symmetries. Although it has been studied on lattice models for very short chains by exhaustive enumeration, it remains a challenge to estimate the designable quantitatively for real proteins. In this study, we designed a new deep neural network model that samples protein sequences given a backbone structure using sequential Monte Carlo method. The sampled sequences with proper weights were used to estimate the designability of several real proteins. The designed sequences were also tested using the latest AlphaFold2 and RoseTTAFold to confirm their foldabilities. We report this as the first study to estimate the designability of real proteins.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Sequence alignment using machine learning for accurate template-based protein structure prediction 97%
- A de novo protein structure prediction by iterative partition sampling, topology adjustment, and residue-level distance deviation optimization 97%
- Scoring Protein Sequence Alignments Using Deep Learning 96%
Similar papers in this journal
Similar papers in this journal
- DISTEMA: distance map-based estimation of single protein model accuracy with attentive 2D convolutional neural network 97%
- Binding affinity prediction for protein-ligand complex using deep attention mechanism based on intermolecular interactions 96%
- Rprot-Vec: A deep learning approach for fast protein structure similarity calculation 96%
Similar papers in this journal
- Pathfinder: protein folding pathway prediction based on conformational sampling 96%
- Hybridized distance- and contact-based hierarchical structure modeling for folding soluble and membrane proteins 95%
- Elucidation of Genome-wide Understudied Proteins targeted by PROTAC-induced degradation using Interpretable Machine Learning 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.