Back

Prediction of physical characteristics of disordered proteins using molecular simulation and physics-informed multiple machine learning strategies

Gonzalez, D. L.; Ibrahim, S.; Atia, G.; Seth, S.; Bhattacharya, A.

2025-06-10 biophysics
10.1101/2025.06.09.658593 bioRxiv
Show abstract

We introduce a novel hybrid machine learning (ML) framework to predict the radius of gyration and other conformational properties of intrinsically disordered proteins (IDPs). Our model integrates sequence information with physical features derived from a coarse-grained (CG) model validated by experimental data. Specifically, we combine hidden states from sequence-based models with 23 physical features projected into a shared latent space, and apply an attention mechanism that assigns weights to each residue to highlight the most informative regions of the sequence. This attention-guided fusion significantly improves predictive accuracy across multiple metrics, including MAPE and MSE, while also enhancing confidence in the predictions. We trained and evaluated our models on Brownian dynamics (BD) simulation results for approximately 7,000 IDPs from the MobiDB database (each with > 99% disorder score). We find that sequence-based models consistently outperform feature-only models, with the GRU achieving the best performance among sequence-only approaches. Moreover, combining sequence and feature information further improves accuracy across all architectures, with the hybrid biGRU model delivering the best overall predictive performance. SHAP analysis reveals the relative importance of physical features, offering model explainability and guiding feature selection. Notably, using a small number of top features often reduces model complexity and improves generalization. Furthermore an integrated gradient analysis reveals that apart from the length of the IDPs, the three parameters (SCD, SHD, and f *) play key role in ML predictions. Our framework provides a fast, interpretable, and scalable tool for predicting IDP behavior, enabling efficient initial screening prior to costly molecular simulations.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.