Back

ProteinCLIP: enhancing protein language models with natural language

Wu, K. E.; Chang, H.; Zou, J.

2024-05-17 bioinformatics
10.1101/2024.05.14.594226 bioRxiv
Show abstract

Language models have enabled a new era of biological sequence modeling. However, extracting meaningful sequence-level embeddings from these models remains challenging. In this work, we introduce ProteinCLIP, which applies contrastive learning between a proteins amino acid sequence and curated text describing its function. ProteinCLIP thus learns to take a pre-trained protein language models sequence embedding and refines it produce a function-centric embedding. We show that this embedding space yields sequence representations that enable state-of-the-art performance across a variety of important yet challenging tasks in the study of proteins - from predicting protein protein interactions to accurately detecting homologous proteins despite low sequence similarity. More broadly, ProteinCLIP demonstrates the effectiveness of multi-modal learning in biological contexts, and how such strategies can help isolate key signals from large models and further improve their utility.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.