Serial Crystallography with Multi-stage Merging of 1000s of Images
Soares, A.; Yamada, Y.; Jakoncic, J.; McSweeney, S.; Sweet, R. M.; Skinner, J.; Foadi, J.; Fuchs, M. R.; Schneider, D. K.; Shi, W.; Andrews, L. C.; Bernstein, H. J.
Show abstract
KAMO and Blend provide particularly effective tools to manage automatically the merging of large numbers of datasets from serial crystallography. The requirement for manual intervention in the process can be reduced by extending Blend to support additional clustering options such as use of more accurate cell distance metrics and use of reflection-intensity correlation coefficients to infer "distances" among sets of reflec- tions. This increases the sensitivity to differences in unit cell parameters and allows for clustering to assemble nearly complete datasets on the basis of intensity or ampli- tude differences. If datasets are already sufficiently complete to permit it, one applies KAMO once and clusters the data using intensities only. If starting from incomplete datasets, one applies KAMO twice, first using cell parameters. In this step we use either the simple cell vector distance of the original Blend, or we use the more sensi- tive NCDist. This step tends to find clusters of sufficient size so that, when merged, each cluster is sufficiently complete to allow reflection intensities or amplitudes to be compared. One then uses KAMO again using the correlation between the reflections having a common hkl to merge clusters in a way sensitive to structural differences that may not have perturbed the cell parameters sufficiently to make meaningful clusters. Many groups have developed effective clustering algorithms that use a measurable physical parameter from each diffraction still or wedge to cluster the data into cate- gories which then can be merged, one hopes, to yield the electron density from a single protein form. Since these physical parameters are often largely independent from one another, it should be possible to greatly improve the efficacy of data clustering software by using a multi-stage partitioning strategy. Here, we have demonstrated one possible approach to multi-stage data clustering. Our strategy is to use unit-cell clustering until merged data is sufficiently complete then to use intensity-based clustering. We have demonstrated that, using this strategy, we are able to accurately cluster datasets from crystals that have subtle differences.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- AHoJ: rapid, tailored search and retrieval of apo and holo protein structures for user-defined ligands 93%
- Local Disordered Region Sampling (LDRS) for Ensemble Modeling of Proteins with Experimentally Undetermined or Low Confidence Prediction Segments 92%
- Ligand Identification in CryoEM and X-ray Maps Using Deep Learning 91%
Similar papers in this journal
- A Robust Method for Collecting X-ray Diffraction Data from Protein Crystals across Physiological Temperatures 96%
- Reciprocalspaceship: A Python Library for Crystallographic Data Analysis 95%
- A drug discovery-oriented non-invasive protocol for protein crystal cryoprotection by dehydration, with application for crystallization screening. 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.