Scaling k-Means for Multi-Million Frames: A Stratified NANI Approach for Large-Scale MD Simulations
Santos, J. B. W.; Chen, L.; Quintana, R. A. M.
Show abstract
We present improved k-means clustering initialization strategies for molecular dynamics (MD) simulations, implemented as part of the N-ary Natural Initiation (NANI) method. Two new deterministic seeding strategies--strat_all and strat_reduced--extend the original NANI approaches and dramatically reduce the clustering runtime while preserving the quality of clustering results. These methods also preserve NANIs reproducible partitioning of well-separated and compact clusters while avoiding the costly iterative seed selection procedures of previous implementations. Testing on the {beta}-heptapeptide and the HP35 systems shows that these new flavors achieved Calinski-Harabasz (CH) and Davies-Bouldin (DB) scores comparable to the previous NANI variant, indicating that the efficiency gains come with no quality decrease. We also show how this new variant can be used to greatly speed up our previously proposed Hierarchical Extended Linkage Method (HELM). These enhancements extend the reach of NANI to accelerate large-scale MD analysis both in stand-alone k-means clustering and as a component of hybrid workflows, and remove a key barrier to routine, scalable, and reproducible exploration of complex conformational ensembles. The improved NANI implementation is accessible through our MDANCE package: https://github.com/mqcomplab/MDANCE.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Interpreting Molecular Dynamics Forces as DeepLearning Gradients Improves Quality Of PredictedProtein Structures 96%
- Molecular insights from conformational ensembles via Machine Learning 95%
- ATOMDANCE: kernel-based denoising and allosteric resonance analysis for functional and evolutionary comparisons of protein dynamics 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.