Contrasting Sequence with Structure: Pre-training Graph Representations with PLMs
Robinson, L. C. B.; Atkinson, T.; Copoiu, L.; Bordes, P.; Pierrot, T.; Barrett, T.
Show abstract
Understanding protein function is vital for drug discovery, disease diagnosis, and protein engineering. While Protein Language Models (PLMs) pre-trained on vast protein sequence datasets have achieved remarkable success, equivalent Protein Structure Models (PSMs) remain underrepresented. We attribute this to the relative lack of high-confidence structural data and suitable pre-training objectives. In this context, we introduce BioCLIP, a contrastive learning framework that pre-trains PSMs by leveraging PLMs, generating meaningful per-residue and per-chain structural representations. When evaluated on tasks such as protein-protein interaction, Gene Ontology annotation, and Enzyme Commission number prediction, BioCLIP-trained PSMs consistently outperform models trained from scratch and further enhance performance when merged with sequence embeddings. Notably, BioCLIP approaches, or exceeds, specialized methods across all benchmarks using its singular pre-trained design. Our work addresses the challenges of obtaining quality structural data and designing self-supervised objectives, setting the stage for more comprehensive models of protein function. Source code is publicly available2.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Pair-EGRET: enhancing the prediction of protein-proteininteraction sites through graph attention networks and protein language models 97%
- ProteinBERT: A universal deep-learning model of protein sequence and function 96%
- Topology-Driven Negative Sampling Enhances Generalizability in Protein-Protein Interaction Prediction 96%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.