Lightway access to AlphaMissense data that demonstrates a balanced performance of this missense mutation predictor
Tordai, H.; Torres, O.; Csepi, M.; Padanyi, R.; Lukacs, G. L.; Hegedus, T.
Show abstract
Single amino acid substitutions can profoundly affect protein folding, dynamics, and function, leading to potential pathological consequences. The ability to discern between benign and pathogenic substitutions is pivotal for therapeutic interventions and research directions. Given the limitations in experimental examination of these variants, AlphaMissense has emerged as a promising predictor of the pathogenicity of single nucleotide polymorphism variants. In our study, we assessed the efficacy of AlphaMissense across several protein groups, such as mitochondrial, housekeeping, transmembrane proteins, and specific proteins like CFTR, using ClinVar data for validation. Our comprehensive evaluation showed that AlphaMissense delivers outstanding performance, with MCC scores predominantly between 0.6 and 0.74. We observed low performance on the CFTR and disordered, membrane-interacting MemMoRF datasets. However, an enhanced performance with CFTR was shown when benchmarked against the CFTR2 database. Our results also emphasize that quality of AlphaFolds predictions can seriously influence AlphaMissense predictions. Most importantly, AlphaMissenses consistent capability in predicting pathogenicity across diverse protein groups, spanning both transmembrane and soluble domains was found. Moreover, the prediction of likely-pathogenic labels for IBS and CFTR coupling helix residues emphasizes AlphaMissenses potential as a tool for pinpointing functionally significant sites. Additionally, to make AlphaMissense predictions more accessible, we have introduced a user-friendly web resource (https://alphamissense.hegelab.org) to enhance the utility of this valuable tool. Our insights into AlphaMissenses capability, along with this online resource, underscore its potential to significantly aid both research and clinical applications.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- HumanMine: advanced data searching, analysis and cross-species comparison. 93%
- PDB NextGen Archive: Centralising Access to Integrated Annotations and Enriched Structural Information by the Worldwide Protein Data Bank 93%
- GenDiS3 database: census on prevalence of protein domain superfamilies of known structure in the entire sequence database 92%
Similar papers in this journal
- CONSTRUCT: an algorithmic tool for identifying functional or structurally important regions in protein tertiary structure 97%
- CATHe: Detection of remote homologues for CATH superfamilies using embeddings from protein language models 94%
- Protein intrinsically disordered regions have a non-random, modular architecture 94%
Similar papers in this journal
- TMSNP: a web server to predict pathogenesis of missense mutations in transmembrane region of membrane proteins 93%
- Predicting Gene Disease Associations With Knowledge Graph Embeddings For Diseases With Curtailed Information 93%
- An Integrative Multitiered Computational Analysis for Better Understanding the Structure and Function of 85 Miniproteins 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.