Context-Aware Protein Representations Using Protein Language Models and Optimal Transport
Patel, S. S.; NaderiAlizadeh, N.
Show abstract
Proteins have different functions in different contexts. As a result, representations that take into account a proteins biological context would allow for a more accurate assessment of its functions and properties. Protein language models (PLMs) generate amino-acid-level (residue-level) embeddings of proteins and are a powerful approach for creating universal protein representations. However, PLMs on their own do not consider context and cannot generate context-specific protein representations. We introduce COPTER, a method that uses optimal transport to pool together a proteins PLM-generated residue-level embeddings using a separate context embedding to create context-aware protein representations. We conceptualize the residue-level embeddings as samples from a probabilistic distribution, and use sliced Wasserstein distances to map these samples against a context-specific reference set, yielding a contextualized protein-level embedding. We evaluate COPTERs performance on three downstream prediction tasks: therapeutic drug target prediction, genetic perturbation response prediction, and TCR-epitope binding prediction. Compared to state-of-the-art baselines, COPTER achieves substantially improved, near-perfect performance in predicting therapeutic targets across cell contexts. It also results in improved performance in predicting responses to genetic perturbations and binding between TCRs and epitopes. The implementation code is available at https://github.com/SahilP113/COPTER.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CellFM: a large-scale foundation model pre-trained on transcriptomics of 100 million human cells 97%
- Designing meaningful continuous representations of T cell receptor sequences with deep generative models 97%
- Sparse Epistatic Regularization of Deep Neural Networks for Inferring Fitness Functions 96%
Similar papers in this journal
Similar papers in this journal
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations 97%
- An Analysis of Protein Language Model Embeddings for Fold Prediction 96%
- scValue: value-based subsampling of large-scale single-cell transcriptomic data for machine and deep learning tasks 95%
Similar papers in this journal
- ProteinBERT: A universal deep-learning model of protein sequence and function 97%
- Topology-Driven Negative Sampling Enhances Generalizability in Protein-Protein Interaction Prediction 97%
- TUGDA: Task uncertainty guided domain adaptation for robust generalization of cancer drug response prediction from in vitro to in vivo settings 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.