Disease-specific variant pathogenicity prediction using multimodal biomedical language models
Liu, Y.; Cooper, D. N.; Yu, H.
Show abstract
Missense variants play a key role in the diagnosis of genetic disorders and in disease risk prediction. Existing methods focus primarily on the prediction of variant effects in terms of their deleteriousness, without taking into account the disease-specific context, and are therefore limited in terms of their utility in real-world diagnosis and decision making. Here, we introduce disease-specific variant pathogenicity prediction (DIVA), a novel deep learning framework that directly predicts specific disease types alongside the probability of deleteriousness for missense variants. Our approach integrates information from two different modalities - protein sequence and disease-related textual annotations - encoded using two pre-trained language models and optimized within a contrastive learning paradigm designed to align variants with relevant diseases in the learned representation space. Our results demonstrate that DIVA outperforms baselines and provides accurate disease predictions with high relevance to clinically curated disease annotations for missense variants. Variant deleteriousness prediction is enhanced by incorporating AlphaMissense scores through learnable weights derived from protein function annotations, which additionally boosts DIVAs ability to accurately classify deleterious variants. Our work provides new insights into variant pathogenicity prediction with awareness of disease specificity, addressing a hitherto unmet need in relation to clinical variant interpretation.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- PreMode predicts mode-of-action of missense variants by deep graph representation learning of protein sequence and structural context 97%
- DeMAG predicts the effects of variants in clinically actionable genes by integrating structural and evolutionary epistatic features 97%
- A probabilistic graphical model for estimating selection coefficient of nonsynonymous variants from human population sequence data 96%
Similar papers in this journal
- Genomics 2 Proteins portal: A resource and discovery tool for linking genetic screening outputs to protein sequences and structures 96%
- Sliding Window INteraction Grammar (SWING): a generalized interaction language model for peptide and protein interactions 96%
- scGPT: Towards Building a Foundation Model for Single-Cell Multi-omics Using Generative AI 95%
Similar papers in this journal
- Learning multi-cellular representations of single-cell transcriptomics data enables characterization of patient-level disease states 94%
- SCOPE: a normalization and copy number estimation method for single-cell DNA sequencing 94%
- Integrative, high-resolution analysis of single cell gene expression across experimental conditions with PARAFAC2-RISE 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.