Back

GermRL: Alleviating The Germline Bias In Autoregressive Antibody Language Models Through Reinforcement Learning

Ludwig, L.; Chungyoun, M.; Gray, J. J.

2026-06-11 bioinformatics
10.64898/2026.06.08.730660 bioRxiv
Show abstract

Antibodies are powerful therapeutics whose antigen specificity arises from sequence diversity shaped during development. Recently, language models trained on large antibody repertoire datasets have enabled the generation and screening of novel candidates, but these models retain a strong germline bias. As AI adoption increases in therapeutic workflows, it is crucial to develop models that harness the diversity of antibodies necessary for the discovery of mutations that encode desirable properties. Previous work explored the germline bias in masked antibody language models, yet the bias in generative autoregressive language models has not yet been addressed. Here, we present GermRL, a lightweight and modular reinforcement learning (RL) framework capable of alleviating the germline bias in pre-trained antibody autoregressive language models through group relative policy optimization (GRPO). GermRL achieves consistent one-shot generation of antibodies that satisfy specified mutation thresholds from germline while maintaining structural plausibility. Under the lowest and highest mutation thresholds tested (5 and 35 mutations from germline), GermRL scores 0.992 and 0.950 pass@1, respectively, compared to 0.398 and 0.034 for the pre-trained language model. Within GermRL, we introduce a key pair of modifications to GRPO that increase training efficiency by discouraging reward hacking under our antibody application. Furthermore, comparison of RL generated and natural antibody sequences reveals how RL based optimization can explore alternative evolutionary mutational patterns and residue compositional strategies while preserving key global properties of natural antibodies, including identifiable germline assignments, embedding-level similarity and comparable developability profiles. Thus, RL-trained generative models optimized to promote antibody mutations through diversity from germline provide a promising framework for navigating the antibody sequence landscape, enabling exploration of novel yet biologically plausible candidates for therapeutic design.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
mAbs
32 papers in training set
Top 0.1%
22.0%
2
Nature Machine Intelligence
70 papers in training set
Top 0.1%
12.7%
3
Journal of Chemical Information and Modeling
238 papers in training set
Top 0.8%
6.8%
4
Briefings in Bioinformatics
354 papers in training set
Top 1%
6.8%
5
PLOS Computational Biology
1863 papers in training set
Top 6%
5.5%
50% of probability mass above
6
Computational and Structural Biotechnology Journal
242 papers in training set
Top 1%
4.1%
7
iScience
1154 papers in training set
Top 5%
3.4%
8
Nature Communications
5641 papers in training set
Top 35%
3.3%
9
Bioinformatics Advances
203 papers in training set
Top 2%
3.2%
10
Scientific Reports
3612 papers in training set
Top 42%
2.5%
11
Cell Systems
201 papers in training set
Top 2%
2.1%
12
ImmunoInformatics
12 papers in training set
Top 0.1%
1.7%
13
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 30%
1.5%
14
Bioinformatics
1204 papers in training set
Top 7%
1.5%
15
Patterns
78 papers in training set
Top 2%
1.3%
16
Communications Chemistry
48 papers in training set
Top 0.8%
1.3%
17
PLOS ONE
5266 papers in training set
Top 55%
1.1%
18
Communications Medicine
113 papers in training set
Top 4%
1.0%
19
Molecular Therapy Nucleic Acids
39 papers in training set
Top 0.8%
1.0%
20
Frontiers in Immunology
638 papers in training set
Top 9%
0.9%
21
Cell Reports Methods
165 papers in training set
Top 4%
0.8%
22
PRX Life
42 papers in training set
Top 0.9%
0.8%
23
Advanced Science
286 papers in training set
Top 9%
0.8%
24
npj Systems Biology and Applications
125 papers in training set
Top 2%
0.8%
25
Nucleic Acids Research
1281 papers in training set
Top 15%
0.6%