Soft Metropolis-Hastings Correction For Generative Model Sampling
Feng, H.; Qiu, P.; Poczos, B.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWMolecular diffusion models suffer from systematic sampling biases that prevent optimal structure formation, resulting in chemically suboptimal molecules with metastable conformations trapped in local energy minima. We introduce Metropolis-Hastings (MH) correction to molecular diffusion models, providing a principled framework to address these systematic sampling biases. The traditional hard accept-reject Metropolis-Hastings corrector creates discontinuous trajectories incompatible with the continuous nature of molecular potential energy surfaces, disrupting proper structure assembly. To address this, we develop a soft Metropolis-Hastings correction that replaces binary acceptance with continuous interpolation weighted by acceptance probabilities, maintaining smooth navigation in the chemical space while providing principled bias correction. We design three molecular-specific variants and demonstrate through extensive experiments on small molecules, drug conformations, and therapeutic antibody CDR-H3 loops that our method consistently improves chemical validity, structural stability, and conformational quality across diverse molecular families. Our method establishes MH correction as a powerful component for molecular generation.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- An Orientationally Averaged Version of the Rotne-Prager-Yamakawa Tensor Provides A Fast But Still Accurate Treatment Of Hydrodynamic Interactions In Brownian Dynamics Simulations Of Biological Macromolecules 92%
- Elucidating Protein Dynamics through the Optimal Annealing of Variational Autoencoders 92%
- Representation of Protein Dynamics Disentangled by Time-structure-based Prior 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.