Unlocking Predictive Power: A Machine Learning Tool Derived from In-Depth Analysis to Forecast the Impact of Missense Variants in Human Filamin C
Nagy, M.; Mlynek, G.; Kostan, J.; Smith, L.; Puehringe, D.; Charron, P.; Rasmussen, T. B.; Bilinska, Z.; Akhtar, M. M.; Syrris, P.; Lopes, L. R.; Elliott, P. M.; Gautel, M.; Carugo, O.; Djinovic-Carugo, K.
Show abstract
Cardiomyopathies, diseases of the heart muscle, are a leading cause of heart failure. An increasing proportion of cardiomyopathies have been associated with specific genetic changes, such as mutations in FLNC, the gene that codes for filamin C. Altogether, more than 300 variants of FLNC have been identified in patients, including a number of single point mutations. However, the role of a significant number of these mutations remains unknown. Here, we conducted a comprehensive analysis, starting from clinical data that led to identification of new pathogenic and non-pathogenic FLNC variants. We selected some of these variants for further characterization that included studies of in vivo effects on the morphology of neonatal cardiomyocytes to establish links to phenotype, and the in vitro thermal stability and structure determination to understand biophysical factors impacting function. We used these findings to compile vast datasets of pathogenic and non-pathogenic variant structures and developed a machine-learning-based neural network (AMIVA-F) to predict the impact of single point mutations. AMIVA-F outperformed most commonly used predictors both in disease related as well as neutral variants, approaching [~]80% accuracy. Taken together, our study documents additional FLNC variants, their biophysical and structural properties, and their link to the disease phenotype. Furthermore, we developed a state-of-the-art web-based server AMIVA-F that can be used for accurate predictions regarding the effect of single point mutations in human filamin C, with broad implications for basic and clinical research.
Matching journals
The top 12 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Structure of transmembrane prolyl 4-hydroxylase reveals unique organization of EF and dioxygenase domains 95%
- Inhibition of cleavage of human complement component C5 and the R885H C5 variant by two distinct high affinity anti-C5 nanobodies 94%
- Conservation of C4BP-binding Sequence Patterns in Streptococcus pyogenes M and Enn Proteins 94%
Similar papers in this journal
- Cardiomyopathy-Associated and Basic Residue Mutations in Myopalladin Alter Actin Binding, Bundling, and Structural Stability 95%
- Identification of 121 variants of honey bee Vitellogenin protein sequences with structural differences at functional sites 94%
- Integrated structural model of the palladin-actin complex using XL-MS, docking, NMR, and SAXS 94%
Similar papers in this journal
- Redefining the architecture of ferlin proteins: insights into multi-domain protein structure and function 93%
- Conserved intramolecular networks in GDAP1 are closely connected to CMT-linked mutations and protein stability 93%
- Neuropathy-related mutations alter the membrane binding properties of the human myelin protein P0 cytoplasmic tail 93%
Similar papers in this journal
- Mutations affecting the N-terminal domains of SHANK3 point to different pathomechanisms in neurodevelopmental disorders. 94%
- Structural consequences of BMPR2 kinase domain mutations causing pulmonary arterial hypertension 93%
- Predicting human and viral protein variants affecting COVID-19 susceptibility and repurposing therapeutics 93%
Similar papers in this journal
- Intracellular helix-loop-helix domain modulates inactivation kinetics of mammalian TRPV5 and TRPV6 channels. 93%
- Identification of ATP2B4 regulatory element containing functional genetic variants associated with severe malaria 93%
- Redefining hypo- and hyper-responding phenotypes of CFTR mutants for understanding and therapy. 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.