Genome-wide prediction of disease variants with a deep protein language model
Brandes, N.; Goldman, G.; Wang, C. H.; Ye, C. J.; Ntranos, V.
Show abstract
Distinguishing between damaging and neutral missense variants is an ongoing challenge in human genetics, with profound implications for clinical diagnosis, genetic studies and protein engineering. Recently, deep-learning models have achieved state-of-the-art performance in classifying variants as pathogenic or benign. However, these models are currently unable to provide predictions over all missense variants, either because of dependency on close protein homologs or due to software limitations. Here we leveraged ESM1b, a 650M-parameter protein language model, to predict the functional impact of human coding variation at scale. To overcome existing technical limitations, we developed a modified ESM1b workflow and functionalized, for the first time, all proteins in the human genome, resulting in predictions for all [~]450M possible missense variant effects. ESM1b was able to distinguish between pathogenic and benign variants across [~]150K variants annotated in ClinVar and HGMD, outperforming existing state-of-the-art methods. ESM1b also exceeded the state of the art at predicting the experimental results of deep mutational scans. We further annotated [~]2M variants across [~]9K alternatively-spliced genes as damaging in certain protein isoforms while neutral in others, demonstrating the importance of considering all isoforms when functionalizing variant effects. The complete catalog of variant effect predictions is available at: https://huggingface.co/spaces/ntranoslab/esm_variants.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Redefining tissue specificity of genetic regulation of gene expression in the presence of allelic heterogeneity 95%
- A Deep Dive into Statistical Modeling of RNA Splicing QTLs Reveals New Variants that Explain Neurodegenerative Disease 94%
- A Scalable Framework for Identifying Allelic Series from Summary Statistics 94%
Similar papers in this journal
- DeMAG predicts the effects of variants in clinically actionable genes by integrating structural and evolutionary epistatic features 95%
- MutPred2: inferring the molecular and phenotypic impact of amino acid variants 95%
- Multi-context genetic modeling of transcriptional regulation resolves novel disease loci 95%
Similar papers in this journal
- The Nucleotide Transformer: Building and Evaluating Robust Foundation Models for Human Genomics 96%
- Genomics 2 Proteins portal: A resource and discovery tool for linking genetic screening outputs to protein sequences and structures 94%
- Biophysics-based protein language models for protein engineering 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.