Back

ProtBLIP2-SST: Protein Function Prediction via BLIP2 with Sequence, Structure, and Text

Chen, Z.; Luo, Q.

2026-07-12 bioinformatics
10.64898/2026.07.10.737551 bioRxiv
Show abstract

Protein function prediction traditionally relies on structured gene ontology (GO) labels or multi-label classifiers. However, these labels or classifiers cannot flexibly describe molecular function, biological process, cellular component, and free-text functional narratives in a single output. In comparison, generation-based approaches offer an intuitive paradigm for flexible free-text protein annotation, with large language models (LLMs) as a representative method for protein-text modeling. Recent efforts on utilizing LLMs for protein semantic understanding and annotation generation have adopted sequence-only encoding or sequence-text contrastive alignment paradigms, yet without explicit consideration of three-dimensional structural information. To address these limitations in current protein function prediction methods, we present ProtBLIP2-SST, a two-stage framework built on the BLIP2 model architecture that bridges protein sequence, structure, and text for open-ended protein functional caption generation. Specifically, we first integrate sequence and structure information through SaProt, a protein language model (PLM) with a structure-aware vocabulary that fuses residue tokens with Foldseek-derived 3Di structural tokens. To empower the LLM to understand protein semantics, we employ a Q-Former (a querying transformer in BLIP2) with learnable query tokens as the cross-modal projector to align protein features from the frozen SaProt encoder and text features from a frozen BiomedBERT via protein-text contrasting, protein-text matching, and protein captioning objectives. After alignment, the protein features are linearly projected and prepended to the prompt embeddings of the LLM for protein captioning fine-tuning with LoRA. Trained on 441k protein-text pairs from Swiss-Prot with corresponding structures from the AlphaFold Database, our ProtBLIP2-SST outperforms sequence-only and sequence-text alignment baselines on protein captioning metrics, with ablation studies demonstrating the effectiveness of integrating structure with sequence information for improved protein understanding. Through a unified two-stage alignment-and-generation pipeline, ProtBLIP2-SST integrates protein sequence and structural information, overcomes the rigidity of traditional GO-centric classification, generating open-ended captions that jointly describe molecular function, subcellular location, and homology context in one single output.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 2%
15.0%
2
Nature Methods
385 papers in training set
Top 0.5%
15.0%
3
Nature Machine Intelligence
70 papers in training set
Top 0.1%
12.4%
4
Nature Communications
5641 papers in training set
Top 24%
6.7%
5
Briefings in Bioinformatics
354 papers in training set
Top 2%
5.1%
50% of probability mass above
6
Cell Systems
201 papers in training set
Top 1%
4.0%
7
Journal of Chemical Information and Modeling
238 papers in training set
Top 1%
3.2%
8
Bioinformatics Advances
203 papers in training set
Top 2%
3.2%
9
Nature Biotechnology
172 papers in training set
Top 2%
2.4%
10
Patterns
78 papers in training set
Top 0.9%
2.4%
11
Journal of Molecular Biology
232 papers in training set
Top 1%
2.1%
12
Nature Computational Science
55 papers in training set
Top 0.5%
1.9%
13
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 28%
1.7%
14
Protein Science
246 papers in training set
Top 2%
1.7%
15
Nucleic Acids Research
1281 papers in training set
Top 9%
1.7%
16
Scientific Reports
3612 papers in training set
Top 56%
1.7%
17
Genome Biology
637 papers in training set
Top 7%
1.1%
18
Advanced Science
286 papers in training set
Top 8%
1.0%
19
PLOS Computational Biology
1863 papers in training set
Top 18%
1.0%
20
Cell Reports Methods
165 papers in training set
Top 3%
1.0%
21
PRX Life
42 papers in training set
Top 0.9%
1.0%
22
Genome Research
468 papers in training set
Top 6%
0.8%
23
NAR Genomics and Bioinformatics
242 papers in training set
Top 4%
0.8%
24
Journal of Computational Biology
48 papers in training set
Top 1%
0.8%
25
iScience
1154 papers in training set
Top 40%
0.6%