Back

ProtNHF: Neural Hamiltonian Flows for Controllable Protein Sequence Generation

Raghavan, B.; Rogers, D. M.

2026-03-06 bioinformatics
10.64898/2026.03.04.709305 bioRxiv
Show abstract

Controllable protein sequence generation remains a central challenge in computational protein design, as most existing approaches rely on retraining, classifier guidance, or architectural modification to impose conditioning. Here we introduce ProtNHF, a generative model that enables continuous, quantitative control over sequence-level properties through analytical bias functions applied exclusively at inference time. ProtNHF builds on neural Hamiltonian flows, where a lightweight Transformer-based potential energy function, inspired by ESM-2, is combined with an explicit kinetic term to define Hamiltonian dynamics in a continuous relaxation of protein sequence space. The model learns a symplectic transport map from a latent Gaussian distribution to protein sequence embeddings via deterministic leapfrog integration, enabling efficient and expressive sampling. In the unconditional setting, generated sequences achieve competitive quality as measured by ESM-2 pseudo-perplexity and AlphaFold2 pLDDT confidence scores. A key advantage of the Hamiltonian formulation is its additive energy structure, which permits external bias potentials to be incorporated directly into the Hamiltonian at inference time without modifying or retraining the learned model. This casts controllable generation in a classical molecular modeling paradigm, where desired properties are enforced by explicit energy shaping. We demonstrate smooth, predictable, and approximately monotonic control over amino acid composition and global properties such as net charge by introducing simple analytical bias terms, including residue-specific chemical potentials and harmonic constraints. The bias strength modulates the values of these properties in generated sequences continuously while preserving structural plausibility and diversity. ProtNHF thus provides a flexible base distribution that can be steered toward different compositional regimes using transparent, physically interpretable energy terms, establishing a general framework for inference-time programmable protein sequence generation.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.