Back

Bayesian multi-state multi-condition modeling of a protein structure based on X-ray crystallography data

Hancock, M.; HOLTON, J. M.; Fraser, J. S.; Adams, P. D.; Sali, A.

2025-10-07 biophysics
10.1101/2025.10.07.680994 bioRxiv
Show abstract

1An atomic structure model of a protein can be computed from a diffraction pattern of its crystal. While most crystallographic studies produce a single set of atomic coordinates, the billions of protein molecules in a crystal sample many conformational modes during data collection. As a result, a "multi-state" model that depicts these conformations could reproduce the X-ray data better than a single conformation, and thus likely be more accurate. Computing such a multistate model is challenging due to a lower data-to-parameter ratio than that for single-state modeling. To address this challenge, additional information could be considered, such as X-ray datasets collected for the same system under distinct experimental conditions (eg, temperature, ligands, mutations, and pressure). Here, we develop, benchmark, and illustrate MultiXray: Bayesian multi-state multi-condition modeling for X-ray crystallography. The input information is several X-ray datasets collected under distinct conditions and a molecular mechanics force field. The model consists of an independent coordinate set for each of several states and the weight of each state under each condition. A Bayesian posterior model density quantifies the match of the model with all X-ray datasets and the force field. A sample of models is drawn from the posterior model density using biased molecular dynamics (MD) simulations. We benchmark MultiXray on simulated CypA X-ray data. Using a second X-ray dataset improves the Rfree from 0.105 to 0.089. We then demonstrate MultiXray on experimental temperature-dependent data for SARS-CoV-2 Mpro. Using multiple X-ray datasets improves Rfree of the PDB-deposited structure from 0.253 to 0.237. MultiXray is implemented in our open-source Integrative Modeling Platform (IMP) software, relying on integration with Phenix, thus making it easily applicable to many studies.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.