PLMView: collaborative protein language model representations for fast and scalable specialized protein function inference
Pho, V.-S.; Bianchi, A. N.; Scarsini, M.; Bowler, C.; Carbone, A.
Show abstract
The functional classification of protein sequences remains a major bottleneck in biology. Although protein language model (PLM)-based approaches have substantially improved broad protein function prediction, most protein sequences still lack precise annotation at the level of specialized functions--the fine-grained molecular roles that define specificity within protein families. We present PLMView, an unsupervised framework for fine-grained protein function classification directly from sequence. PLMView reframes protein function inference as a relational problem: instead of embedding sequences in isolation, it positions them within a collaborative functional space defined by comparisons with PLM embeddings of anchor sequences, thereby capturing subtle sequence-function relationships. Without requiring labeled data, family-specific training, or PLM fine-tuning, PLMView accurately distinguishes specialized functions among homologous proteins and highlights residues likely to determine functional specificity. The method achieves high precision while remaining computationally efficient, classifying approximately 10,000 sequences with 1,000 anchors in under 40 minutes; compared with pooled-embedding approaches and, in challenging cases, Sequence Similarity Networks, PLMView provides finer and more biologically coherent functional resolution, while achieving more than 10-fold speed-up over SSN reconstruction on datasets of this scale. Applications to thioredoxins, visual opsins, and Tara Oceans environmental diatom cold-shock proteins show that PLMView can move from interpretable residue-level determinants in well-studied protein families to large-scale environmental functional discovery, linking molecular specialization to ecological distribution and transcriptional deployment across the global ocean.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- EvoRMD: Integrating Biological Context and Evolutionary RNA Language Models for Interpretable Prediction of RNA Modifications 94%
- EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations 93%
- Deciphering the Sequence Basis and Application of Transcriptional Initiation Regulation in Plant Genomes Through Deep Learning 93%
Similar papers in this journal
- RegFormer: A Single-Cell Foundation Model Powered by Gene Regulatory Hierarchies 94%
- DGAT: A Dual-Graph Attention Network for Inferring Spatial Protein Landscapes from Transcriptomics 94%
- Constructing Ensemble Gene Functional Networks Capturing Tissue/condition-specific Co-expression from Unlabled Transcriptomic Data with TEA-GCN 94%
Similar papers in this journal
- Ancestral Reconstruction of Protein Interaction Networks 95%
- Medusa: software to build and analyze ensembles of genome-scale metabolic network reconstructions 93%
- Engineering indel and substitution variants of diverse and ancient enzymes using Graphical Representation of Ancestral Sequence Predictions (GRASP) 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.