Back

ProtAug: An Empirical Investigation of pLM-Guided Data Augmentation for Protein Sequence Prediction Tasks

Chen, Z.; Wang, R.; Luo, Q.

2026-07-11 bioinformatics
10.64898/2026.07.10.737545 bioRxiv
Show abstract

Protein language models (pLMs) offer great potential for protein sequence analysis, yet the scarcity of labeled data often limits their effectiveness in fine-tuning. Data augmentation is a promising remedy, but systematic evaluation of augmentation strategies for protein sequences remains limited, and the conditions under which augmentation confers downstream benefits are not well understood. In this paper, we systematically investigate pLM-guided substitution-based augmentation across seven protein prediction tasks. We propose ProtAug, a framework that leverages encoder-based (ESM-2) and autoregressive (ProtGPT2) pLMs to generate augmented sequences with user-controlled variation levels. Our investigation focuses on four questions: (Q1) whether pLM-synthesized sequences preserve more original signals than simpler methods, (Q2) to what extent augmentation improves prediction performance, (Q3) how variation levels affect downstream accuracy across tasks and models, and (Q4) whether biological plausibility is a necessary condition for achieving improvement. Our experimental results show that: (1) ProtAug Esm generally preserves motifs and structural similarity better than simple substitution, often comparable to homology retrieval; (2) augmentation yields consistent but task-dependent improvements, with ProtAug Esm achieving the best or second-best performance in 5 out of 7 tasks at 10% variation; (3) low-to-moderate variation levels (2-30%) perform best overall, although high-variation augmentation can benefit certain structure-related tasks; (4) the necessity of biological plausibility is task- and variation-dependent--while semantic preservation correlates with performance at low-to-moderate variation levels, improved generalization at high variation levels suggests that regularization effects, rather than label preservation, can also drive performance gains.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 1%
18.5%
2
Briefings in Bioinformatics
354 papers in training set
Top 0.6%
9.7%
3
Bioinformatics Advances
203 papers in training set
Top 0.4%
7.9%
4
Journal of Chemical Information and Modeling
238 papers in training set
Top 0.8%
6.7%
5
BMC Bioinformatics
457 papers in training set
Top 2%
4.8%
6
Frontiers in Bioinformatics
49 papers in training set
Top 0.1%
4.3%
50% of probability mass above
7
Journal of Computational Biology
48 papers in training set
Top 0.2%
4.3%
8
PLOS Computational Biology
1863 papers in training set
Top 9%
4.0%
9
BioData Mining
22 papers in training set
Top 0.1%
3.5%
10
NAR Genomics and Bioinformatics
242 papers in training set
Top 1%
3.2%
11
Journal of Cheminformatics
29 papers in training set
Top 0.3%
2.4%
12
Scientific Reports
3612 papers in training set
Top 44%
2.4%
13
Computational and Structural Biotechnology Journal
242 papers in training set
Top 3%
2.1%
14
PLOS ONE
5266 papers in training set
Top 47%
1.9%
15
Protein Science
246 papers in training set
Top 2%
1.7%
16
Genome Biology
637 papers in training set
Top 6%
1.4%
17
Journal of Molecular Biology
232 papers in training set
Top 3%
1.1%
18
Cell Systems
201 papers in training set
Top 3%
1.1%
19
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 35%
1.1%
20
Nature Machine Intelligence
70 papers in training set
Top 2%
1.1%
21
GigaScience
212 papers in training set
Top 4%
1.0%
22
Nucleic Acids Research
1281 papers in training set
Top 12%
1.0%
23
Nature Communications
5641 papers in training set
Top 57%
0.8%
24
Nature Methods
385 papers in training set
Top 6%
0.8%
25
Genomics, Proteomics & Bioinformatics
16 papers in training set
Top 0.2%
0.8%
26
IEEE/ACM Transactions on Computational Biology and Bioinformatics
38 papers in training set
Top 1%
0.8%
27
Journal of Chemical Theory and Computation
140 papers in training set
Top 1%
0.8%
28
Biology Methods and Protocols
61 papers in training set
Top 3%
0.6%