Interpreting Protein Language Models: high attention sites predict functional regions
Pribus, S. J.; Altman, R. B.; Nayar, G.
Show abstract
Computational proteomics has revolutionized biomedical research, guiding targeted experimental exploration to accelerate protein-based mechanistic discovery. Protein Language Models (PLMs) enable scalable, resource-efficient study; through large-scale training on only primary protein sequences, PLMs generate vector representations of protein structure that have been shown to capture biochemical, evolutionary, and structural properties. A core component of PLMs is the attention mechanism, which specifically captures long-range interactions across a protein sequence in attention matrices. Using the Evolutionary Scale Modelling 2 (ESM-2) PLM, we previously developed a novel method to identify "High Attention" (HA) sites. HA sites are specific residues that ESM-2 assigns the most attention to early during encoding. Here, we further characterize these HA sites across structural and functional metrics. Using unsupervised clustering, we find HA sites can be categorized as "structural core", "structural pathogenic", "core pathogenic", or "low-confidence". We further use AlphaMissense pathogenicity predictions and the pan-cancer analysis of whole genomes (PCAWG)-labeled pathogenic variant positions to show that HA sites predict protein regions with high pathogenic risk. Finally, we explore the utility of HA sites for suggesting candidate binding sites, identifying multiple cancer protein examples where HA sites identified regions with previously undiscovered high interaction likelihood and thus potential therapeutic utility. Our work demonstrates the biological interpretability of PLM representations and offers a valuable method to prioritize functionally relevant protein residues for targeted biomedical research.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Protein prediction models support widespread post-transcriptional regulation of protein abundance by interacting partners 94%
- DART-ID increases single-cell proteome coverage 93%
- A semi-supervised Bayesian approach for simultaneous protein sub-cellular localisation assignment and novelty detection 93%
Similar papers in this journal
- FAVA: High-quality functional association networks inferred from scRNA-seq and proteomics data 93%
- Proteome-level assessment of origin, prevalence and function of Leucine-Aspartic Acid (LD) motifs 93%
- PEPerMINT: Peptide Abundance Imputation in Mass Spectrometry-based Proteomics using Graph Neural Networks 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.