Back

Accelerated sampling of protein dynamics using BioEmu augmented molecular simulation

Bhakat, S.

2026-01-07 biophysics
10.64898/2026.01.07.698041 bioRxiv
Show abstract

We introduce a workflow that integrates BioEmu-generated conformational ensemble with physics-based molecular simulations and Markov State Models to sample Boltzmann-weighted conformational populations across biomolecules. Molecular simulations initiated from BioEmu ensemble capture active-to-inactive transitions in CDK2 and BRAF, two members of the serine-threonine kinase family, and elucidate how the disease-causing V600E mutation in BRAF drives population shifts among distinct metastable states relative to the wild type. Furthermore, we combined BioEmu ensemble with experimental cryo-EM data to construct all-atom conformational ensembles of biomolecules. In comparison to the AlphaFold2 reduced multiple sequence alignment (rMSA-AF2) approach, BioEmu-generated ensembles sample a broader conformational space for serine-threonine kinases but fail to capture conformational heterogeneity in several cases, including Glycine transporter 1 (GlyT1), a membrane transporter, and plasmepsin-II (PlmII), an aspartic protease. Systems where side-chain conformational heterogeneity governs protein dynamics such as cryptic pocket opening in PlmII or transitions between multiple metastable states in GlyT1; molecular simulations initiated from BioEmu generated ensemble do not capture the full spectrum of conformational heterogeneity. Overall, this study presents a straightforward framework for integrating generative AI based protein emulators with statistical physics to recover Boltzmann-weighted conformational ensembles at scale, while also highlighting critical limitations that necessitate careful, system-specific analysis when interpreting protein conformational landscapes.

Published in Journal of Chemical Information and Modeling (predicted rank #2) · training set

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.