EpiPred: A gene-specific machine learning model for classifying missense variants in the epilepsy-related gene STXBP1
Calhoun, J.; Wang, C.; Biar, C.; Gunti, J.; Lee, J.; Geller, A.; Hong, J.; Schnell, S.; Dang, L.; Wang, Y.; Parent, J.; Isom, L.; Uhler, M.; Mefford, H.; Ross, M.; Pulido, V.; Carvill, G.
Show abstract
Missense variants in the STXBP1 gene are a frequent cause of early-onset developmental and epileptic encephalopathies and related neurodevelopmental disorders, but the clinical interpretation of these variants remains a major challenge. Most reported STXBP1 missense variants are classified as variants of uncertain significance (VUS), complicating diagnosis, counseling, and patient eligibility for precision therapies. Here, we developed EpiPred, a gene-specific machine learning classifier that predicts the pathogenicity of STXBP1 missense variants by integrating computational features with empirical evidence from cellular assays. Trained on a curated set of pathogenic and benign variants, EpiPred outperformed leading global prediction tools in accuracy, sensitivity, and specificity. We validated the models predictions using functional assays that measure protein abundance, solubility, stability, and interaction with the SNARE complex partner syntaxin 1. These biochemical readouts aligned closely with model outputs and enabled reclassification of several likely misdiagnosed variants. We deployed EpiPred in an interactive web application that allows clinicians, researchers, and patients to explore predictions for all possible missense variants in STXBP1. Our approach illustrates the power of gene-specific predictive modeling combined with experimental validation to improve variant interpretation and diagnostic resolution. By identifying likely pathogenic STXBP1 variants, including those that may respond to emerging therapies such as molecular chaperones, EpiPred supports more precise genetic diagnoses and offers a generalizable framework for other clinically relevant genes in neurological disease. ONE SENTENCE SUMMARYEpiPred improves STXBP1 variant interpretation, enabling precision genetic diagnoses and promoting access to targeted precision therapies
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Stretch-activated ion channel TMEM63B associates with developmental and epileptic encephalopathies and progressive neurodegeneration 95%
- Multi-parametric analysis of 58 SYNGAP1 variants reveal impacts on GTPase signaling, localization and protein stability 94%
- Loss of C2orf69 defines a fatal auto-inflammatory mitochondriopathy in Humans and Zebrafish 94%
Similar papers in this journal
- Genomic analyses of glycine decarboxylase neurogenic mutations yield a large scale prediction model for prenatal disease. 95%
- Missense variants causing Wiedemann-Steiner syndrome preferentially occur in the KMT2A-CXXC domain and are accurately classified using AlphaFold2 94%
- Genome mining yields new disease-associated ROMK variants with distinct defects 94%
Similar papers in this journal
- Long-read genome sequencing for the diagnosis of neurodevelopmental disorders 94%
- Functional characterization of pathogenic SATB2 missense variants identifies distinct effects on chromatin binding and transcriptional activity 93%
- Identification and validation of novel candidate risk genes in endocytic vesicular trafficking associated with esophageal atresia and tracheoesophageal fistulas 93%
Similar papers in this journal
- Systematic analysis of genetic and phenotypic characteristics reveals antisense oligonucleotide therapy potential for one-third of neurodevelopmental disorders 95%
- Genome-wide prediction of pathogenic gain- and loss-of-function variants from ensemble learning of diverse feature set 94%
- The impact of damaging epilepsy and cardiac genetic variant burden in sudden death in the young 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.