Back

Interpretable Machine Learning Uncovers Structural Determinants of Wnt-Wls Binding Specificity from Extended Atomistic Simulations

Callahan, T. J.; Shi, J.; Cheng, K. J.; Sauer, M. A.; Pogorelov, T. V.; Capponi, S.

2025-08-22 biochemistry
10.1101/2025.08.18.670971 bioRxiv
Show abstract

The Wnt protein family plays a critical role in cell development, with each Wnt protein interacting differently with the Wls membrane protein through distinct binding residues. A direct comparison and elucidation of the molecular mechanisms underlying Wnt-Wls binding across the diverse Wnt family remain challenging, owing to variations in sequence length and amino acid composition among Wnt proteins, which can affect their binding affinity and trafficking efficiency via Wls. Here we combine extended atomistic molecular dynamics simulations with supervised machine learning to elucidate binding specificity among four Wnt proteins, selected based on experimental structure availability and scientific relevance. We implement a local structure alignment algorithm to enable cross-system matching and comparison of residue interactions, and we apply a two-stage clustering strategy to reduce feature redundancy and facilitate robust feature selection. After training a Random Forest classifier that achieved high predicting accuracy, our feature importance analysis reveals both previously known and novel key residue pairs responsible for distinguishing among the Wnt systems. Our findings highlight that the binding specificity across different systems arises from the distributed nature of interactions across the protein binding surface and demonstrate how interpretable machine learning can effectively uncover crucial biophysical interactions. Importantly, our integrated strategy is generalizable to other systems and provides a data-driven approach for analyzing protein- protein interactions and guiding experimental validation or therapeutic targeting.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.