Enhancing Missense Variant Classification in Predicted Intrinsically Disordered Regions
Gnanaolivu, R.; Hart, S.
Show abstract
The accurate classification of missense variants is a fundamental challenge in genomics, particularly for those within intrinsically disordered regions (IDRs) where the performance of existing computational predictors is suboptimal. To address this, we developed a machine learning model that extends traditional missense tools with properties that infer globular IDR conformation, phase separation, and protein embeddings. Using ClinVar variant classifications as ground truth, AlphaMissense, EVE, and ESM1b were the highest scoring unsupervised in silico missense predictors for IDR variants. Our baseline model, using only IDR-specific features achieved competitive performance on the hold-out test set with a PR-AUC of 0.800. Critically, when these IDR features were combined with these methods we saw significant overall improvement. The AlphaMissense-Enhanced model increased its PR-AUC from 0.807 to 0.931. Similarly, ESM1b-Enhanced improved PR-AUC from 0.679 to 0.878 and EVE increased from 0.591 to 0.918. These results demonstrate the effectiveness of our enhancements for classifying missense variants in IDRs and highlight its ability to complement existing in silico missense predictors.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- 3Cnet: Pathogenicity prediction of human variants using knowledge transfer with deep recurrent neural networks 93%
- Beyond the Leaderboard: Leveraging Predictive Modeling for Protein-Ligand Insights and Discovery 93%
- acmgscaler: An R package and Colab for standardised gene-level variant effect score calibration within the ACMG/AMP framework 93%
Similar papers in this journal
- LYRUS: A Machine Learning Model for Predicting the Pathogenicity of Missense Variants 95%
- CanDrivR-CS: A Cancer-Specific Machine Learning Framework for Distinguishing Recurrent and Rare Variants 93%
- Estimating Protein Complex Model Accuracy Using Graph Transformers and Pairwise Similarity Graphs 93%
Similar papers in this journal
- When splicing is not all or none: Implications for variant classification 93%
- IMPROVE-DD: Integrating Multiple Phenotype Resources Optimises Variant Evaluation in genetically determined Developmental Disorders 93%
- Characteristics predicting reduced penetrance variants in the high-risk cancer predisposition gene TP53 92%
Similar papers in this journal
- Large scale analyses of genotype-phenotype relationships of glycine decarboxylase mutations and neurological disease severity. 94%
- MENDELSEEK: An algorithm that predicts Mendelian Genes and elucidates what makes them special 93%
- Mutation severity spectrum of rare alleles in the human genome is predictive of disease type 93%
Similar papers in this journal
- The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients 94%
- ProtAlign-ARG: Antibiotic Resistance Gene Characterization Integrating Protein Language Models and Alignment-Based Scoring 93%
- Protein Embeddings Predict Binding Residues in Disordered Regions 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.