Back

Journal of Chemical Theory and Computation

American Chemical Society (ACS)

All preprints, ranked by how well they match Journal of Chemical Theory and Computation's content profile, based on 140 papers previously published here. The average preprint has a 0.08% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
OpenCafeMol with 3SPN.2 DNA model: GPU Acceleration for Long-Time Coarse-Grained Chromatin Simulations

Yamauchi, M.; Murata, Y.; Niina, T.; Takada, S.

2026-03-19 biophysics 10.64898/2026.03.18.712524 medRxiv
Top 0.1%
73.5%
Show abstract

There is a growing demand for molecular dynamics simulations to explore longer timescale behavior of giant protein-DNA complexes such as chromatin. To address this need, we extended OpenCafeMol, a GPU-accelerated residue-level coarse-grained molecular dynamics simulator originally developed for proteins and lipids, to support 3SPN.2 and 3SPN.2C DNA models. We also implemented a hydrogen-bond-type many-body potential to model DNA-protein interactions more accurately. To further improve computational efficiency, we introduced a localized scheme for calculating base-pairing and cross-stacking interactions. Benchmark tests show that OpenCafeMol on a single GPU achieves up to 200-fold speed-up for DNA-only systems and up to 100-fold speed-up for DNA-protein complexes compared to CPU-based simulations. To demonstrate the capability of our implementation for long-timescale biological processes, we simulated an archaeal SMC-ScpA complex undergoing DNA translocation via segment capture (a proposed mechanism for DNA loop extrusion) in the presence of a DNA-bound obstacle. We observed continuous captured-loop growth accompanied by obstacle bypass within the segment capture framework.

2
Proteus software for physics-based protein design

Mignon, D.; Druart, K.; Opuu, V.; Polydorides, S.; Villa, F.; Gaillard, T.; Michael, E.; Archontis, G.; Simonson, T.

2020-07-01 biophysics 10.1101/2020.06.30.179549 medRxiv
Top 0.1%
65.7%
Show abstract

We describe methods and software for physics-based protein design. The folded state energy combines molecular mechanics with Generalized Born solvent. Sequence and conformation space are sampled with Replica Exchange Monte Carlo, assuming one or a few fixed protein backbone structures and discrete side chain rotamers. Whole protein design and enzyme design are presented as illustrations. Full redesign of three PDZ domains was done using a simple, empirical, unfolded state model. Designed sequences were very similar to natural ones. Enzyme redesign exploited a powerful, adaptive, importance sampling approach that allows the design to directly target substrate binding, reaction rate, catalytic efficiency, or the specificity of these properties. Redesign of tyrosyl-tRNA synthetase stereospecificity is reported as an example.Competing Interest StatementThe authors have declared no competing interest.View Full Text

3
CGRig: a rigid-body protein model with residue-level interaction sites for long-time and large-scale protein assembly simulation

Teshirogi, Y.; Terada, T.

2026-03-24 biophysics 10.64898/2026.03.21.713350 medRxiv
Top 0.1%
64.9%
Show abstract

Molecular dynamics (MD) simulations are a powerful tool for investigating biomolecular dynamics underlying biological functions. However, the accessible spatiotemporal scales of conventional all-atom simulations remain limited by high computational costs. Coarse-graining reduces these costs by decreasing the number of interaction sites and enabling longer timesteps. In extreme cases, proteins are represented as single spherical particles; while such approximations facilitate cellular-scale simulations, they often sacrifice essential structural information, such as molecular shape and interaction anisotropy. Here, we present CGRig, a rigid-body protein model with residue-level interaction sites designed for long-time, large-scale simulations. In CGRig, each protein is treated as a single rigid-body embedding residue-level interaction sites. Its translational and rotational motions are described by the overdamped Langevin equation incorporating a shape-dependent friction matrix. Intermolecular interactions are calculated using G[o]-like native contact potentials, Debye-Huckel electrostatics, and volume exclusion. We validated that CGRig accurately reproduces the translational and rotational diffusion coefficients expected from the friction matrix for an isolated protein. For dimeric systems, the model successfully maintained native complex structures. Furthermore, two initially separated proteins converged into the correct complex with an association rate consistent with all-atom simulations. Notably, CGRig achieved a simulation performance exceeding 17 s/day for a 1,024-molecule system. These results demonstrate that CGRig provides an efficient framework for simulating protein assembly while retaining residue-level interaction specificity, making it a valuable tool for investigating large-scale biomolecular self-assembly.

4
Rare-Event Sampling using a Reinforcement Learning-Based Weighted Ensemble Method

Yang, D.; Goldberg, A.; Chong, L.

2024-10-11 biophysics 10.1101/2024.10.09.617475 medRxiv
Top 0.1%
60.7%
Show abstract

Despite the power of path sampling strategies in enabling simulations of rare events, such strategies have not reached their full potential. A common challenge that remains is the identification of a progress coordinate that captures the slow relevant motions of a rare event. Here we have developed a weighted ensemble (WE) path sampling strategy that exploits reinforcement learning to automatically identify an effective progress coordinate among a set of potential coordinates during a simulation. We apply our WE strategy with reinforcement learning to three benchmark systems: (i) an egg carton-shaped toy potential, (ii) an S-shaped toy potential, and (iii) a dimer of the HIV-1 capsid protein (C-terminal domain). To enable rapid testing of the latter system at the atomic level, we employed discrete-state synthetic molecular dynamics trajectories using a generative, fine-grained Markov state model that was based on extensive conventional simulations. Our results demonstrate that using concepts from reinforcement learning with a weighted ensemble of trajectories automatically identifies relevant progress coordinates among multiple candidates at a given time during a simulation. Due to the rigorous weighting of trajectories, the simulations maintain rigorous kinetics.

5
NEAT-DNA: A Chemically Accurate, Sequence-Dependent Coarse-Grained Model for Large-Scale DNA Simulations

Riveros, I.; Zhang, B.

2025-11-08 biophysics 10.1101/2025.11.07.687145 medRxiv
Top 0.1%
60.4%
Show abstract

DNAs physical properties play a central role in genome organization and regulation, but simulating its behavior across biologically relevant scales remains a major computational challenge. Coarse grained DNA models have enabled faster simulations, yet they often sacrifice chemical accuracy or produce unphysical conformations, limiting their utility for studying genome structure. A key difficulty has been constructing a model that is both efficient enough for large-scale simulations and faithful to the molecular mechanics of DNA. Here we introduce NEAT-DNA, a new coarse-grained DNA model that resolves longstanding limitations in physical realism and parameter optimization. By combining a physically principled energy formulation with a unified training framework that integrates data from both atomistic simulations and experiments, NEAT-DNA accurately reproduces sequence-dependent structure and flexibility while remaining computationally efficient. This approach marks a significant advance over previous models, which either lacked sequence specificity or introduced distortions inconsistent with experimental observations. NEAT-DNA bridges this gap, offering a high-fidelity yet tractable representation of DNA suitable for exploring chromatin folding. More broadly, it provides a foundation for large-scale simulations that couple molecular detail with gene-level chromatin organization, opening new avenues for predictive modeling in structural genomics.

6
Efficient RNA Folding Simulation via a Structure-Based Single-Site-Per-Nucleotide Model

Thornton, T.; Lin, X.

2025-12-14 biophysics 10.64898/2025.12.13.694107 medRxiv
Top 0.1%
60.3%
Show abstract

Computational modeling of large RNA structures and their dynamics is essential for uncovering the molecular mechanisms underlying various genomic processes and RNA-regulated cellular functions. Residue-resolution modeling is an effective approach for simulating large biomolecular structures while preserving essential sequence and struc-tural features presented in atomic structures. Here, we implemented a structure-based single-site-per-nucleotide (SSPN) RNA model using the GPU-accelerated OpenMM 1 software and evaluated its computational efficiency and accuracy by simulating RNA hairpins unfolding under force. Our simulations compare favorably with an earlier, more detailed RNA model and quantitatively reproduce experimentally measured ther-modynamic properties of RNA under mechanical stretching. This SSPN model enables scalable and accurate simulations of large RNA ensembles, such as long non-coding RNAs and RNA liquid-liquid phase separation.

7
DiffPIE: Guiding Deep Generative Models to Explore Protein Conformations under External Interactions

Wang, Y.; Chen, M.

2025-04-30 biophysics 10.1101/2025.04.27.650875 medRxiv
Top 0.1%
59.5%
Show abstract

In recent years, many foundation generative models have been developed to pre-dict structures of molecules and materials. Although these foundation models have achieved great success, it is challenging to collect enough data to train foundation generative models. One such example is to predict protein conformations with protein-environmental interactions (PEI), such as interactions introduced by organic linkers or material surfaces. We propose a physics-guided route to extrapolate foundation mod-els beyond their training domain. Our method couples a pretrained deep generative model with explicit, physics-based interaction potentials for PEI, steering sampling to-ward conformations consistent with external constraints without any retraining or fine-tuning. We demonstrate accurate and efficient conformation prediction of (i) cyclic peptide with organic linkers and (ii) peptide adsorbed on the gold surface. The gen-erated structures serve as high-quality initial conditions for downstream simulations, providing a general, systematic approach to extend foundation models to proteins under system-specific environmental interactions.

8
Light Martini water accelerates sampling in coarse-grained molecular dynamics simulations

Elgendy, A.; Zeipelt, A. P.; Schäfer, L. V.

2026-08-03 biophysics 10.64898/2026.08.03.741232 medRxiv
Top 0.1%
57.1%
Show abstract

Molecular dynamics (MD) simulations of slow biomolecular processes, such as exploration of the conformational ensembles of intrinsically disordered proteins (IDPs), are computationally demanding. Although coarse-grained (CG) models can substantially speed up the simulations compared to all-atom MD, the sampling challenge can still be significant for large systems and long time scales. Here, we present light Martini water, a low-viscosity water model that accelerates sampling in MD simulations with the Martini CG force field. We systematically reduced the mass of the Martini water beads and verified stable, accurate integration of the equations of motion with 20 fs time steps, as typically used in Martini simulations. Light Martini water has a reduced mass of 20 amu (compared to 72 amu in the standard water model), yielding up to a 2.68-fold increase in the sampling rate of IDP chain reconfiguration in water and a 16 % increase in the lateral diffusion of lipids in a POPC bilayer. Equilibrium properties remained unaffected by the mass scaling, and the speedup was achieved without compromising simulation accuracy. The water model is trivial to implement, has no computational overhead, and should be universally applicable to Martini simulations.

9
Fine-tuning molecular mechanics force fields to experimental free energy measurements

Rufa, D.; Fass, J.; Chodera, J. D.

2025-01-08 biophysics 10.1101/2025.01.06.631610 medRxiv
Top 0.1%
56.3%
Show abstract

Alchemical free energy methods using molecular mechanics (MM) force fields are essential tools for predicting thermodynamic properties of small molecules, especially via free energy calculations that can estimate quantities relevant for drug discovery such as affinities, selectivities, the impact of target mutations, and ADMET properties. While traditional MM forcefields rely on hand-crafted, discrete atom types and parameters, modern approaches based on graph neural networks (GNNs) learn continuous embedding vectors that represent chemical environments from which MM parameters can be generated. Excitingly, GNN parameterization approaches provide a fully end-to-end differentiable model that offers the possibility of systematically improving these models using experimental data. In this study, we treat a pretrained GNN force field--here, espaloma-0.3.2--as a foundation simulation model and fine-tune its charge model using limited quantities of experimental hydration free energy data, with the goal of assessing the degree to which this can systematically improve the prediction of other related free energies. We demonstrate that a highly efficient "one-shot fine-tuning" method using an exponential (Zwanzig) reweighting free energy estimator can improve prediction accuracy without the need to resimulate molecular configurations. To achieve this "one-shot" improvement, we demonstrate the importance of using effective sample size (ESS) regularization strategies to retain good overlap between initial and fine-tuned force fields. Moreover, we show that leveraging low-rank projections of embedding vectors can achieve comparable accuracy improvements as higher-dimensional approaches in a variety of data-size regimes. Our results demonstrate that linearly-perturbative fine-tuning of foundation model electrostatic parameters to limited experimental data offers a cost-effective strategy that achieves state-of-the-art performance in predicting hydration free energies on the FreeSolv dataset.

10
Simple Adjustment of Intra-nucleotide Base-phosphate Interaction in OL3 AMBER Force Field Improves RNA Simulations

Mlynsky, V.; Kuhrova, P.; Stadlbauer, P.; Krepl, M.; Otyepka, M.; Banas, P.; Sponer, J.

2023-09-06 biophysics 10.1101/2023.09.05.556403 medRxiv
Top 0.1%
56.2%
Show abstract

Molecular dynamics (MD) simulations represent an established tool to study RNA molecules. Outcome of MD studies depends, however, on the quality of the used force field (ff). Here we suggest a correction for the widely used AMBER OL3 ff by adding a simple adjustment of nonbonded parameters. The reparameterization of Lennard-Jones potential for the -H8...O5- and -H6...O5- atom pairs addresses an intra-nucleotide steric clash occurring in the type 0 base-phosphate interaction (0BPh). The non-bonded fix (NBfix) modification of 0BPh interactions (the NBfix0BPh modification) was tuned via reweighting approach and, subsequently, tested using extensive set of standard and enhanced sampling simulations of both unstructured and folded RNA motifs. The modification corrects minor but visible intra-nucleotide clash for the anti nucleobase conformation. We observed that structural ensembles of small RNA benchmark motifs simulated with the NBfix0BPh modification provide better agreement with experiments. No side-effects of the modification were observed in standard simulations of larger structured RNA motifs. We suggest that the combination of OL3 RNA ff and NBfix0BPh modification is a viable option to improve RNA MD simulations.

11
Environment-conditioned design of alpha-helical peptides

Conde-Torres, D.; Garcia-Fandino, R.; Pineiro, A.

2026-05-08 biophysics 10.64898/2026.05.07.723485 medRxiv
Top 0.1%
56.0%
Show abstract

Designing peptide sequences that remain stable and selective across heterogeneous environments remains a central challenge in biomolecular modeling. Here we introduce an interpretable, physics-based Hamiltonian for environment-conditioned design of -helical peptide sequences. The model integrates helix propensities, pairwise interactions, electrostatics, anisotropic solvent exposure, and interfacial geometry into a unified energy function. To enable comparison across sequence lengths and environments, all contributions are rescaled and expressed as Z-scores relative to random sequence ensembles, yielding a normalized design landscape with balanced physical terms. This formulation defines a structured optimization problem that can be explored using exact, heuristic, and hybrid quantum- classical approaches without modification of the underlying model. The Hamiltonian recovers polar and apolar limits, discriminates experimentally characterized water-soluble and transmembrane -helical peptide sequences, and captures the preferential stabilization of membrane-active sequences at anionic interfaces over non-functional controls. It further enables multi-objective and selective design, generating candidate sequences with tunable environmental specificity.

12
From gHBfix to NBfix: Reweighting-Driven Refinement of Hydrogen-Bond Interactions in RNA Force Fields

Mlynsky, V.; Kuehrova, P.; Bussi, G.; Otyepka, M.; Sponer, J.; Banas, P.

2026-03-21 biophysics 10.64898/2026.03.20.713292 medRxiv
Top 0.1%
56.0%
Show abstract

Understanding RNA structural dynamics is essential for elucidating its biological functions, and molecular dynamics (MD) simulations provide an important atomistic complement to experimental approaches. However, the predictive power of MD is fundamentally limited by the accuracy of the underlying empirical Force Fields (FFs), particularly in capturing the delicate balance of non-bonded interactions. Here, we present a systematic reparameterization strategy that replaces the external gHBfix19 hydrogen-bond (H-bond) correction potential with an equivalent set of NBfix Lennard-Jones modifications within a state-of-the-art RNA FF. Using a quantitatively converged temperature replica-exchange MD ensemble of the GAGA tetraloop, we employed a reweighting-based optimization protocol to derive NBfix parameters that reproduce the thermodynamic effects of the original gHBfix19 terms. Sequential optimization of individual gHBfix19 components proved essential to ensure stable and transferable parameter refinement. The resulting fully reformulated NBfix-based variant, termed OL3CP-NBfix19, was validated on a representative set of RNA motifs, including tetranucleotides, A-form duplexes, and tetraloops. Across all tested systems, its performance is comparable to that of the reference gHBfix19 FF. By embedding the H-bond corrections directly into the standard non-bonded framework, the NBfix formulation eliminates external biasing potentials, simplifies practical deployment, and reduces computational overhead. Beyond this specific reparameterization, our results demonstrate a practical workflow for translating targeted H-bond corrections into native FF terms for efficient biomolecular simulations.

13
Artificial Intelligence Boosted Molecular Dynamics

Do, H. N.; Miao, Y.

2023-03-27 bioinformatics Community evaluation 10.1101/2023.03.25.534210 medRxiv
Top 0.1%
55.8%
Show abstract

We have developed a new Deep Boosted Molecular Dynamics (DBMD) method. Probabilistic Bayesian neural network models were implemented to construct boost potentials that exhibit Gaussian distribution with minimized anharmonicity, thereby allowing for accurate energetic reweighting and enhanced sampling of molecular simulations. DBMD was demonstrated on model systems of alanine dipeptide and the fast-folding protein and RNA structures. For alanine dipeptide, 30ns DMBD simulations captured up to 83-125 times more backbone dihedral transitions than 1{micro}s conventional molecular dynamics (cMD) simulations and were able to accurately reproduce the original free energy profiles. Moreover, DBMD sampled multiple folding and unfolding events within 300ns simulations of the chignolin model protein and identified low-energy conformational states comparable to previous simulation findings. Finally, DBMD captured a general folding pathway of three hairpin RNAs with the GCAA, GAAA, and UUCG tetraloops. Based on Deep Learning neural network, DBMD provides a powerful and generally applicable approach to boosting biomolecular simulations. DBMD is available with open source in OpenMM at https://github.com/MiaoLab20/DBMD/.

14
Exploring Conformational Transitions of RNA Dimers via Machine Learning Potentials

Medrano Sandonas, L.; Tolmos Nehme, M.; Cofas-Vargas, L. F.; Olivos-Ramirez, G. E.; Cuniberti, G.; Poblete, S.; Poma, A. B.

2026-02-26 biophysics 10.64898/2026.02.25.707885 medRxiv
Top 0.1%
55.4%
Show abstract

RNA is a flexible biopolymer that adopts diverse conformations while forming structural motifs essential for its function. Classical RNA force fields often show limited transferability and inefficient sampling of transitions between stable states, particularly in moderately large RNA. To address these limitations, quantum-informed machine learning (ML) potentials have recently emerged as a promising alternative, offering improved accuracy and transferability relative to classical force fields. Here, we assess ML potentials for exploring RNA conformations using the adenine-adenine dinucleoside monophosphate (ApA) dimer, a fundamental RNA building block. We generated an extensive quantum-mechanical (QM) dataset for ApA conformations obtained from temperature replica exchange molecular dynamics (TREMD) simulations. Despite its small size, the ApA dimer exhibits six conformations in which quantum effects and solvent-mediated interactions play crucial roles. Using this dataset, we parameterized ML potentials based on the equivariant MACE architecture and informed by both ab-initio and semi-empirical data. The resulting potentials reproduce key conformational features of the ApA system, including base stacking, sugar puckering, and backbone flexibility, and provide broader coverage of structural transitions than the general-purpose SO3LR and MACE-OFF24 models. These findings highlight the importance of quantum-accurate RNA force fields towards the structural and energetic characterization of RNA complexes.

15
An essential dynamics-based elastic network model to unravel the conformational dynamics of DNA, RNA, and protein-nucleic acid complexes

Cannariato, M.; Scaramozzino, D.; Lee, B. H.; Deriu, M. A.; Orellana, L.

2026-03-13 biophysics 10.64898/2026.03.11.710985 medRxiv
Top 0.1%
55.2%
Show abstract

The flexibility of DNA and RNA is known to play a central role in numerous biological processes, including chromatin organization and gene regulation. While a wide range of computational approaches have been developed to investigate the conformational dynamics and flexibility of proteins, analogous methods for nucleic acids remain comparatively underexplored. Elastic Network Models (ENMs) - coarse-grained mechanical representations in which macromolecules are modeled as networks of nodes connected by elastic springs - have been successfully applied to proteins, often allowing to capture experimentally observed conformational changes through a small number of harmonic normal modes. Building on a previously validated three-bead ENM for RNA, here we introduce edENM, an essential dynamics-refined ENM for DNA, RNA, and protein-nucleic acid complexes, parametrized using a diverse set of Molecular Dynamics simulations. The vibrational modes of the new edENM show good agreement with NMR data and experimental ensembles, while avoiding the unrealistic and localized deformability of previous ENM parametrizations. Additionally, we integrated this new edENM into eBDIMS, a Brownian Dynamics-based framework that enables the simulation of large-scale and anharmonic conformational transitions in protein assemblies. In this way, we are now able to explore functional motions in large protein-nucleic acid complexes such as chromatin subunits and ribosomes.

16
Binding Affinity Estimation From Restrained Umbrella Sampling Simulations

Govind Kumar, V.; Agrawal, S.; Suresh Kumar, T. K.; Moradi, M.

2021-10-28 biophysics 10.1101/2021.10.28.466324 medRxiv
Top 0.1%
55.1%
Show abstract

The protein-ligand binding affinity quantifies the binding strength between a protein and its ligand. Computer modeling and simulations can be used to estimate the binding affinity or binding free energy using data- or physics-driven methods or a combination thereof. Here, we discuss a purely physics-based sampling approach based on biased molecular dynamics (MD) simulations, which in spirit is similar to the stratification strategy suggested previously by Woo and Roux. The proposed methodology uses umbrella sampling (US) simulations with additional restraints based on collective variables such as the orientation of the ligand. The novel extension of this strategy presented here uses a simplified and more general scheme that can be easily tailored for any system of interest. We estimate the binding affinity of human fibroblast growth factor 1 (hFGF1) to heparin hexasaccharide based on the available crystal structure of the complex as the initial model and four different variations of the proposed method to compare against the experimentally determined binding affinity obtained from isothermal calorimetry (ITC) experiments. Our results indicate that enhanced sampling methods that sample along the ligand-protein distance without restraining other degrees of freedom do not perform as well as those with additional restraint. In particular, restraining the orientation of the ligands plays a crucial role in reaching a reasonable estimate for binding affinity. The general framework presented here provides a flexible scheme for designing practical binding free energy estimation methods.

17
Towards Convergence in Folding Simulations of RNA Tetraloops: Comparison of Enhanced Sampling Techniques and Effects of Force Field Corrections

Mlynsky, V.; Janecek, M.; Kuhrova, P.; Frohlking, T.; Otyepka, M.; Bussi, G.; Banas, P.; Sponer, J.

2021-12-01 biophysics 10.1101/2021.11.30.470631 medRxiv
Top 0.1%
54.6%
Show abstract

Atomistic molecular dynamics (MD) simulations represent established technique for investigation of RNA structural dynamics. Despite continuous development, contemporary RNA simulations still suffer from suboptimal accuracy of empirical potentials (force fields, ffs) and sampling limitations. Development of efficient enhanced sampling techniques is important for two reasons. First, they allow to overcome the sampling limitations and, second, they can be used to quantify ff imbalances provided they reach a sufficient convergence. Here, we study two RNA tetraloops (TLs), namely the GAGA and UUCG motifs. We perform extensive folding simulations and calculate folding free energies ({Delta}Gfold) with the aim to compare different enhanced sampling techniques and to test several modifications of the nonbonded terms extending the AMBER OL3 RNA ff. We demonstrate that replica exchange solute tempering (REST2) simulations with 12-16 replicas do not show any sign of convergence even when extended to time scale of 120 s per replica. However, combination of REST2 with well-tempered metadynamics (ST-MetaD) achieves good convergence on a time-scale of 5-10 s per replica, improving the sampling efficiency by at least two orders of magnitude. Effects of ff modifications on {Delta}Gfold energies were initially explored by the reweighting approach and then validated by new simulations. We tested several manually-prepared variants of gHBfix potential which improve stability of the native state of both TLs by up to ~2 kcal/mol. This is sufficient to conveniently stabilize the folded GAGA TL while the UUCG TL still remains under-stabilized. Appropriate adjustment of van der Waals parameters for C-H...O5 base-phosphate interaction are also shown to be capable of further stabilizing the native states of both TLs by ~0.6 kcal/mol.

18
Exploring RNA conformational ensembles in silico: progress and challenges

Roeder, K.; Stirnemann, G.; Meuret, L.; Barquero-Morera, D.; Forget, S.; Wales, D. J.; Pasquali, S.

2026-02-18 molecular biology 10.64898/2026.02.18.706514 medRxiv
Top 0.1%
54.3%
Show abstract

RNA function is intrinsically linked to its structural polymorphism, with molecules exploring the heterogeneous conformational ensembles resulting from complex energy landscapes. These landscapes arise from competing interactions, small energetic separations between microstates, and strong coupling to the environment, posing significant challenges for both experimental characterization and molecular simulation. In this chapter, we review current computational strategies that aim to explore RNA conformational ensembles in silico, with a specific focus on energy landscape-based approaches and atomistic simulations. We discuss key limitations related to sampling efficiency, force-field accuracy, and ensemble analysis, and illustrate their impact through case studies on a self-cleaving ribozyme and an H-type pseudoknot. Finally, we highlight emerging directions, including closer integration with experimental data and the growing role of machine learning, which will probably reinforce the predictive power of in silico RNA energy landscape exploration.

19
Multi-Agent Reinforcement Learning-based Adaptive Sampling for Conformational Sampling of Proteins

Kleiman, D. E.; Shukla, D.

2022-05-31 biophysics 10.1101/2022.05.31.494208 medRxiv
Top 0.1%
52.8%
Show abstract

Machine Learning is increasingly applied to improve the efficiency and accuracy of Molecular Dynamics (MD) simulations. Although the growth of distributed computer clusters has allowed researchers to obtain higher amounts of data, unbiased MD simulations have difficulty sampling rare states, even under massively parallel adaptive sampling schemes. To address this issue, several algorithms inspired by reinforcement learning (RL) have arisen to promote exploration of the slow collective variables (CVs) of complex systems. Nonetheless, most of these algorithms are not well-suited to leverage the information gained by simultaneously sampling a system from different initial states (e.g., a protein in different conformations associated with distinct functional states). To fill this gap, we propose two algorithms inspired by multi-agent RL that extend the functionality of closely-related techniques (REAP and TSLC) to situations where the sampling can be accelerated by learning from different regions of the energy landscape through coordinated agents. Essentially, the algorithms work by remembering which agent discovered each conformation and sharing this information with others at the action-space discretization step. A stakes function is introduced to modulate how different agents sense rewards from discovered states of the system. The consequences are threefold: (i) agents learn to prioritize CVs using only relevant data, (ii) redundant exploration is reduced, and (iii) agents that obtain higher stakes are assigned more actions. We compare our algorithm with other adaptive sampling techniques (Least Counts, REAP, TSLC, and AdaptiveBandit) to show and rationalize the gain in performance.

20
Toward Accurate RNA Folding Thermodynamics: Evaluation of Enhanced Sampling Methods for Force Field Benchmarking

Kuehrova, P.; Mlynsky, V.; Otyepka, M.; Sponer, J.; Banas, P.

2026-01-21 biophysics 10.64898/2026.01.19.700441 medRxiv
Top 0.1%
52.8%
Show abstract

Biologically functional RNAs operate near marginal stability, and their rugged free-energy landscapes and profound structural dynamics - typically not captured by structural biology experiments - play decisive roles. Atomistic molecular dynamics (MD) simulations provide a unique means to characterize these features. However, the applicability of atomistic MD is currently limited by accessible simulation timescales and, most importantly, by force-field (FF) accuracy. Folding free energies ({Delta}G{degrees}fold) of small RNA motifs represent well-defined targets for quantitative benchmarking of RNA FFs. In practice, however, obtaining thermodynamic estimates that are sufficiently robust for direct comparison with experimental data remains highly challenging, even for small RNA systems, and many published studies rely on sampling that is not fully converged. Here, we systematically assess the performance of widely used advanced enhanced-sampling techniques using the 8-mer r(gcGAGAgc) tetraloop as a representative benchmark system. We test temperature replica exchange (T-REMD), two solute-tempering variants of replica exchange (REST2 and REHT), as well as well-tempered metadynamics and on-the-fly probability enhanced sampling combined with solute tempering (ST-MetaD and ST-OPES). Among the tested approaches, T-REMD proves to be the most robust, yielding reproducible folding equilibria and consistent estimates of {Delta}G{degrees}fold after approximately 20 s of simulation time, independent of the initial folded or unfolded conformational ensemble. Our results provide practical guidelines for selecting sampling protocols suitable for quantitative RNA benchmarks and lay the foundation for systematic validation and future refinement of RNA FFs.