Back

Fast prediction of protein flexibility

Praznikar, J.

2025-12-03 bioinformatics
10.64898/2025.11.30.691417 bioRxiv
Show abstract

Advances in hardware have made molecular dynamics (MD) simulations of protein structures faster and more accessible to the scientific community. However, accurately estimating protein flexibility using MD remains computationally demanding, especially for large systems and long time scales. Several MD-based resources - including MdMD, the DynamD database, and more recently ATLAS and mdCATH - now provide MD trajectories for thousands of proteins, enabling the development of predictive models. Here, the Graphlet Degree Vector (GDV) is introduced as a lightweight, fast, and easy-to-implement linear model for predicting protein flexibility directly from atom coordinates. GDV is a 15-dimensional feature vector that captures local packing and the spatial connectivity of each atom with its nearby neighbors. Trained on a subset of globular-like proteins from the ATLAS database, the GDV model achieves a Spearman correlation of 0.828 compared to MD data. The model trained on ATLAS dataset was further evaluated on independent Nuclear Magnetic Resonance and cryo-electron microscopy datasets, demonstrating the robustness and generalizability of the GDV-based approach. A key advantage of the GDV model is that it requires no additional external or experimental data and can be applied in near real time (on the order of 10 seconds) even for large proteins with 20,000 atoms on a standard desktop or laptop. Overall, the results show that a lightweight, fast, and purely coordinate-based model can provide accurate and generalizable predictions of protein flexibility across diverse folds and sizes. The source code is available in the GitHub repository https://github.com/jure-praznikar/FastProtFlex. The data required for model training are available at https://doi.org/10.5281/zenodo.17771418.

Published in Bioinformatics (predicted rank #1) · training set

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.