TooTranslator: Zero-Shot Classification of Specific Substrates for Transport Proteins by Language Embedding Alignment of Proteins and Chemicals
Ataei, S.; Butler, G.
Show abstract
Transmembrane transport proteins mediate selective movement of ions and metabolites across membranes. Experimental characterization of their substrate specificity is limited. For novel class discovery with limited data, zero-shot learning aims to assign labels to test samples whose label has not been seen previously during training of the model. We introduce TooTranslator, a regression-based model that aligns embeddings from pre-trained protein (ProtBERT), chemical (ChemBERTa), and text (SciB-ERT) language models into a shared latent space, enabling substrate prediction by minimizing distances between protein and substrate embeddings. TooTranslator tackles zero-shot learning for the task of predicting the specific substrate of transmembrane transport proteins. Models using four loss functions are evaluated on protein test sets with seen and unseen substrates. We find no statistically significant difference in performance of the four loss functions. For tests with seen substrates, the models compare with the state-of-the art in performance. For tests with unseen (but known) substrate a top-k approach is needed as only 4.3% of test cases predict the correct label as the nearest label. For top-10 that rises to 17%; for top-50 rises to 54%, and top-100 reaches 80%. TooTranslator demonstrates the potential of multimodal embedding alignment for open-world protein function inference.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- FAPM: Functional Annotation of Proteins using Multi-Modal Models Beyond Structural Modeling 97%
- Combining protein sequences and structures with transformers and equivariant graph neural networks to predict protein function 96%
- Pair-EGRET: enhancing the prediction of protein-proteininteraction sites through graph attention networks and protein language models 96%
Similar papers in this journal
- MULAN: Multimodal Protein Language Model for Sequence and Structure Encoding 96%
- KSMoFinder - Knowledge graph embedding of proteins and motifs for predicting kinases of human phosphosites 96%
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 96%
Similar papers in this journal
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 96%
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 96%
- COMIC: Explainable Drug Repurposing via Contrastive Masking for Interpretable Connections 95%
Similar papers in this journal
- CASTER-DTA: Equivariant Graph Neural Networks for Predicting Drug-Target Affinity 96%
- PRIEST - Predicting viral mutations with immune escape capability of SARS-CoV-2 using temporal evolutionary information 96%
- EGRET: Edge Aggregated Graph Attention Networks and Transfer Learning Improve Protein-Protein Interaction Site Prediction 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.