Multilingual model improves zero-shot predictionof disease effects on proteins
Chen, R.; Palpant, N.; Foley, G.; Boden, M.
Show abstract
Predicting the functional impact of genetic variants remains a fundamental challenge in genomics. Existing models focus on protein-intrinsic defects yet overlook regulatory constraints embedded within coding sequences. Here, we couple a codon language model (CaLM) with a protein language model (ESM-2) to dissect the drivers of variant pathogenicity. On ClinVar data, both modalities contribute near-equally to distinguishing pathogenic from benign variants. Evaluation across Deep Mutational Scanning and CRISPR-Based Genome Editing platforms in ClinMAVE reveals that loss-of-function variants are governed primarily by residue-level features, whereas gain-of-function variants show a greater relative contribution from codon-level constraints, albeit in a gene-specific manner. A controlled comparison of identical variants in BRCA1 and TP53 further suggests that codon-level signals are elevated in the endogenous genomic context. Together, these findings indicate that pathogenicity reflects both the "product'' and the "process,'' and that the experimental platform may influence which dimension is observable.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- MTSplice predicts effects of genetic variants on tissue-specific splicing 95%
- Variant effect predictor correlation with functional assays is reflective of clinical classification performance 95%
- DelSIEVE: cell phylogeny model of single nucleotide variants and deletions from single-cell DNA sequencing data 94%
Similar papers in this journal
Similar papers in this journal
- DeMAG predicts the effects of variants in clinically actionable genes by integrating structural and evolutionary epistatic features 96%
- MutPred2: inferring the molecular and phenotypic impact of amino acid variants 95%
- PreMode predicts mode-of-action of missense variants by deep graph representation learning of protein sequence and structural context 95%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.