Validating folding energy estimates as a method for variant interpretation
Elwes, C.; Alcraft, R.; Lister, H.; Smith, P. A.; Shorthouse, D.; Hall, B. A.
Show abstract
Interpretation of variants of uncertain significance remains a major problem in genomic analysis. Whilst statistical models can be used to predict pathogenicity, they offer no insights into the biophysical mechanism of variant action, and genomic data available for training is biased towards the subpopulations who have access. Protein misfolding has been found to act as a frequent mechanism for loss of gene or domain activity, where it is typically responsible for [~]2/3 of disease-causing variants and somatic mutations. The accuracy of energy predictions however has consistently been challenged by highly variable correlation coefficients reported from different proteins, and the unknown impact of alternative structures where available. Here we address this directly through a systematic analysis of mega-scale folding experimental results, enabled by a fully automated predictive pipeline based on FoldX. We find that whilst absolute correlation coefficients are mediocre for three highly studied proteins (0.30-0.31), the correlation coefficient alone does not capture the full predictive power of the estimates. Specifically, we find a clear linear relationship between experimental and theoretical result, with a small number of outlier residues responsible for reducing the correlation. We show that the quantitative accuracy of predictions can be improved by aggregating estimates taken from different structures, and that the problematic outlier residues can be both empirically and theoretically identified, allowing us to flag low-confidence values. Our findings not only provide a framework for identifying problematic mutations in advance but also offers new insights into potential improvements of the FoldX protocol for more accurate protein stability predictions. Our insights support the use of FoldX in computational saturation screens to support variant analysis.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Modeling Alternate Conformations with Alphafold2 via Modification of the Multiple Sequence Alignment 95%
- Improved protein complex prediction with AlphaFold-multimer by denoising the MSA profile 94%
- Predicting changes in protein thermodynamic stability upon point mutation with deep 3D convolutional neural networks 94%
Similar papers in this journal
Similar papers in this journal
- Dynamic Allostery Highlights the Evolutionary Differences between the CoV-1 and CoV-2 Main Proteases 95%
- Modeling Coronavirus Spike Protein Dynamics: Implications for Immunogenicity and Immune Escape 95%
- Interpreting Molecular Dynamics Forces as DeepLearning Gradients Improves Quality Of PredictedProtein Structures 94%
Similar papers in this journal
- Deep learning enables the atomic structure determination of the Fanconi Anemia core complex from cryoEM 92%
- Structures of the germline-specific Deadhead and Thioredoxin T proteins from Drosophila melanogaster reveal unique features among Thioredoxins 92%
- Structural insights on the substrate-binding proteins of the Mycobacterium tuberculosis mammalian-cell-entry (Mce) 1 and 4 complexes 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.