Generating Structurally Diverse Therapeutic Peptides with GFlowNet
Wijaya, E.
Show abstract
Reinforcement learning approaches for therapeutic peptide generation suffer from mode collapse, converging to narrow regions of sequence space even when explicit diversity penalties are applied. Fine-grained analysis reveals persistent mode-seeking behavior invisible to standard diversity metrics. We propose GFlowNet for peptide generation, which samples sequences proportionally to reward rather than maximizing expected reward. This objective provides diversity through proportional sampling without requiring explicit output diversity penalties. Comparing against GRPO with explicit diversity enforcement, GFlowNet achieves substantially more uniform sequence sampling and fewer repetitive motifs. Critically, when diversity mechanisms are removed from the reward, GRPO collapses completely while GFlowNet maintains natural diversity. These results demonstrate that proportional sampling is inherently robust to reward function design, offering a key advantage for drug discovery pipelines requiring diverse candidates.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- PLMFit : Benchmarking Transfer Learning with Protein Language Models for Protein Engineering 95%
- CASTER-DTA: Equivariant Graph Neural Networks for Predicting Drug-Target Affinity 95%
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.