Limitations of the refolding pipeline for de novo protein design
Korbeld, K. T.; Viliuga, V.; Fürst, M. J. L. J.
Show abstract
With the emergence of powerful deep learning-based tools, computational protein design has become a widely accessible technique. Nowadays, it is possible to perform both sequence and structure design in a matter of minutes, making the technology attractive to the broader scientific community. In protein design campaigns, one of the most common in silico strategies to evaluate how well a sequence encodes a target structure is the so-called self-consistency or refolding pipeline. In this approach, a structure prediction model is used to refold the designed sequence to probe whether it is compatible with the intended structure, and is evaluated via two metrics linked to experimental success: the confidence score of the predicted structure (pLDDT) and the self-consistency root-mean-square deviation (scRMSD), which measures how closely the refolded structure matches the target. In this work, we systematically evaluate how different models and structure prediction settings impact these metrics, and to what extent they can be used to reliably filter sequence design candidates. We show that evolutionary information can obscure folding models abilities to assess sequence-structure compatibility, reducing the predictive performance of refolding metrics for experimental success, particularly for designs that share homology with natural sequences. We further highlight limitations of refolding metrics, including their sensitivity to structural features, such as flexibility. Our findings raise awareness of potential pitfalls in refolding-based evaluation and support more informed use of these metrics in protein design campaigns.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Deep-Learning Structure Elucidation from Single-Mutant Deep Mutational Scanning 98%
- Improved protein structure refinement guided by deep learning based accuracy estimation 97%
- Hierarchical design of multi-scale protein complexes by combinatorial assembly of oligomeric helical bundle and repeat protein building blocks 97%
Similar papers in this journal
- Neural Network-Derived Potts Models for Structure-Based Protein Design using Backbone Atomic Coordinates and Tertiary Motifs 98%
- COLLAPSE: A representation learning framework for identification and characterization of protein structural sites 97%
- Evaluation of AlphaFold Antibody-Antigen Modeling with Implications for Improving Predictive Accuracy 96%
Similar papers in this journal
- Predicting structures of large protein assemblies using combinatorial assembly algorithm and AlphaFold2 97%
- US-align: Universal Structure Alignments of Proteins, Nucleic Acids, and Macromolecular Complexes 96%
- OpenFold: Retraining AlphaFold2 yields new insights into its learning mechanisms and capacity for generalization 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.