DeltaMut: An Integrative Database of AlphaFold2-Derived Missense Variant Structures
Qorri, E.; Adam, K.; Takacs, B.; Shemesh, S.; Buzafalvi, D.; Varga, V.; Pekker, E.; Pinter, L.; Hegedüs, Z.; Csanyi, B.; Haracska, L.
Show abstract
The widespread use of next-generation sequencing has led to a surge in the number of identified variants with uncertain effects on protein function. These variants pose a significant challenge in diagnostics and hinder patient treatment strategies. Numerous variant effect predictors (VEPs) are available to assess variant impact, but they primarily rely on sequence-derived information. The recent development of AlphaFold2 has raised questions about whether information retrieved from wild-type or predicted structures of missense variants can improve the predictive power of these algorithms. While the AlphaFold Protein Structure Database serves as a valuable resource for wild-type protein structures, a large-scale collection of missense variant structures is not available, limiting current efforts to wild-type conformations and a handful of modeled variants. To address this limitation, we developed DeltaMut, a comprehensive database containing over 77,000 protein structures, including 65,000 pathogenic and neutral missense variants. All structural models were generated using ParaFold, a high-performance computing-optimized implementation of AlphaFold2. The large-scale and systematic generation of variant protein structures distinguish DeltaMut as a unique resource for both expansive statistical studies and detailed, case-specific investigations of variant-induced structural changes. Furthermore, the DeltaMut database is freely accessible without registration. HighlightsO_LIDeltaMut is currently the largest database of AlphaFold2-predicted variant structures. C_LIO_LIContains 77,713 structures covering 12,101 wild-type and 65,612 variant proteins. C_LIO_LI70.6% of predicted structures have high or very high confidence (pLDDT [≥] 70). C_LIO_LIFreely accessible web server with visualization and download of variant models. C_LI
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- A3D Database: Structure-based Protein Aggregation Predictions for the Human Proteome 96%
- Protlego: A Python package for the analysis and design of chimeric proteins 95%
- Local Disordered Region Sampling (LDRS) for Ensemble Modeling of Proteins with Experimentally Undetermined or Low Confidence Prediction Segments 95%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- A new web resource to predict the impact of missense variants at protein interfaces using 3D structural data: Missense3D-PPI 93%
- ModelCIF: An extension of PDBx/mmCIF data representation for computed structure models 93%
- Integrating multimeric threading with high-throughput experiments for structural interactome of Escherichia coli 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.