High-resolution dissection of concept acquisition in different families of protein language models
Whitfield, S. T.; Marty, T.; Vernon, R. M.; Langmead, C. J.; Sridhar, D.; Fournier, Q.
Show abstract
Protein language models have been increasingly successful on tasks ranging from fitness prediction to functional design, yet what biological knowledge they acquire and where it is encoded within their internal representations remain underexplored. Through a high-resolution layer-by-layer interpretability analysis of 8 models from the ESM2 and AMPLIFY families on 22 concepts from human proteome annotations, we found that these models encode concepts of increasing levels of complexity along their depth: basic physicochemical properties and linear motifs are best captured by early-layer embeddings, secondary structure from subsequent layers, and domain-level semantics from middle layers. Principal component projections of these embeddings showed that they separate biologically meaningful protein groupings, and molecular-biology-inspired interventions demonstrated that pLM embeddings can discriminate phosphomimic-active from inactive mutants. Perhaps surprisingly, we observed that pretraining data and compute had a greater impact on the linear emergence of biological concepts than scaling up parameters. By revealing where biological knowledge is captured in pLMs and which choices shape its emergence, our work offers insights to develop more robust, biologically grounded protein language models.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CaLMPhosKAN: Prediction of General Phosphorylation Sites in Proteins via Fusion of Codon Aware Embeddings with Amino Acid Aware Embeddings and Wavelet-based Kolmogorov Arnold Network 96%
- Sitetack: A Deep Learning Model that Improves PTM Predictionby Using Known PTMs 96%
- ProteinBERT: A universal deep-learning model of protein sequence and function 96%
Similar papers in this journal
- Artificial neural networks enable genome-scale simulations of intracellular signaling 95%
- Structure-Based Function Prediction using Graph Convolutional Networks 95%
- PreMode predicts mode-of-action of missense variants by deep graph representation learning of protein sequence and structural context 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.