Back

Integrating Millions of Years of Evolutionary Information into Protein Structure Models for Function Prediction

Ma, R.; He, C.; Zhang, Z.; Zheng, H.; Duan, L.

2025-11-13 bioinformatics
10.1101/2025.11.11.687944 bioRxiv
Show abstract

BackgroundUnderstanding life processes relies on accurate protein function prediction, fundamentally requiring the integration of evolutionary information encoded in sequences with spatial characteristics from 3D structures. Existing approaches often face limitations, however, by either over-relying on sequence, using simplified structural representations instead of fine-grained spatial details, or failing to capture the synergistic relationship between sequence and structure, compounded by challenges in acquiring annotated data. ResultsTo address these issues, we propose a novel contrast-aware pre-training framework, ESMSCOP. ESMSCOP leverages a state-of-the-art protein language model to harness evolutionary insights embedded in sequences, and introduces a new encoder to fuse topological and fine-grained spatial structural features. By employing a contrastive pre-training strategy with auxiliary supervision, ESMSCOP effectively bridges the sequence-structure gap, yielding rich and informative representations. ConclusionsExtensive experiments conducted on multiple benchmark datasets demonstrate that ESMSCOP achieves superior performance in protein function prediction tasks compared to existing methods. Furthermore, it shows strong performance even when utilizing relatively less pre-training data compared to some large-scale models.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.