Which pLM to choose?
Senoner, T.; Koludarov, I.; Guenther, J.; Shehu, A.; Rost, B.; Bromberg, Y.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWProtein-language models (pLMs) provide a novel means for mapping the protein space. Which of these new maps best advances specific biological analyses, however, is not obvious. To elucidate the principles of model selection, we benchmarked fourteen pLMs, spanning several orders of magnitude in parameter count, across a hundred million protein pairs, to assess how well they capture sequence, structure, and function similarity. For each model, we distinguish inherent information, i.e. signal recoverable from raw-embedding distances, and extractable information, i.e. signal revealed through additional supervised training. Three key results emerge. First, pLM protein representation space is inherently different from the space of biological protein representations, i.e. sequences or structures. Here, a size-performance paradox is salient - mid-scale foundation models are as good as much larger ones in reflecting all tested biological properties. Second, pLM representations compress and store biological information in proportion to model size. That is, a lightweight feed-forward network can be trained on embedding pairs to predict said biological properties well - a capacity dividend. Finally, we observe that a task-specific learning radically reshapes the embedding space, gaining inherent understanding of the task, but garbling any further extractions. In other words, smaller pLMs can provide efficient and compute-light general insight. Larger models are advantageous only when fine-tuning is planned to accomplish a specific task. Furthermore, representations generated by "specialist" models are not immediately generalizable throughout protein biology. Thus, for pLMs, bigger isnt always better.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Unsupervised protein embeddings outperform hand-crafted sequence and structure features at predicting molecular function 97%
- Effect of Tokenization on Transformers for Biological Sequences 96%
- DistilProtBert: A distilled protein language model used to distinguish between real proteins and their randomly shuffled counterparts 96%
Similar papers in this journal
- Joint representation of molecular networks from multiple species improves gene classification 94%
- Controllable Protein Design via Autoregressive Direct Coupling Analysis Conditioned on Principal Components 93%
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 93%
Similar papers in this journal
- COLLAPSE: A representation learning framework for identification and characterization of protein structural sites 93%
- MAHOMES II: A webserver for predicting if a metal binding site is enzymatic 93%
- Neural Network-Derived Potts Models for Structure-Based Protein Design using Backbone Atomic Coordinates and Tertiary Motifs 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.