ETAP-CLF: an ESM3-based transformer attention framework for binary protein classification
Ren, J.; Jiang, H.; Li, P.; Yang, X.; Mei, L.; Tong, H.; Lin, L.
Show abstract
Binary protein classification supports diverse tasks in computational biology, including pathway-membership inference and sequence-based candidate prioritization. Protein language models generate information-rich residue-level representations, but downstream classifiers commonly compress them using fixed pooling operations that may discard task-relevant sequence context. We present ETAP-CLF, a compact framework that combines pretrained per-residue ESM3 embeddings with lightweight transformer contextualization and learned attention pooling to classify variable-length proteins and generate residue-level attention scores. The ESM3 parameters remained frozen, and the same ETAP-CLF architecture and hyperparameter configuration were used across ferroptosis-, senescence-, and pyroptosis-associated protein prediction. ETAP-CLF achieved AUROCs of 0.98, 0.95 and 0.91 for these tasks, respectively. In the ferroptosis benchmark, ETAP-CLF outperformed the evaluated published models. These results demonstrate that a common downstream design can adapt to multiple process-associated classification tasks without fine-tuning the task-specific model architecture. ETAP-CLF provides a generalizable approach for sequence-based protein prioritization and a basis for broader evaluation across binary protein-classification problems.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Hybrid Deep Learning with Protein Language Models and Dual-Path Architecture for Predicting IDP Functions 95%
- Cracking the black box of deep sequence-based protein-protein interaction prediction 95%
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.