Context-Aware Hydrophobicity Modeling: HydroMap and FastHydroMap
Lobo, S.; Najafi, S.; Shea, J.-E.; Shell, M. S.
Show abstract
Hydrophobicity governs a vast range of phenomena, from protein-protein interactions to nanomaterial assembly, and can be rigorously quantified by the dewetting free energy (Fdewet) of a molecule or surface. However, hydrophobicity remains widely treated as an additive property of amino acid identity, obscuring the fact that waters response is a collective property of the surface, shaped by curvature, chemical patterning, and neighboring residues. Direct calculation of Fdewet via specialized molecular simulations captures this collective behavior but is prohibitively slow, leaving in place broadly-used, decades-old sequence-based hydropathy scales that neglect the physics of solvation. Here we show that residue-level Fdewet can be predicted from a compact set of local water features (water structural signatures and residue-water potential energy) extracted from a brief and inexpensive all-atom simulation. We embed this insight in two models: HydroMap, which predicts Fdewet directly from water features, and FastHydroMap, a computationally inexpensive graph neural network surrogate trained on HydroMap that requires no solvent simulation. HydroMap and FastHydroMap capture context-dependent hydrophobicity that classical, sequence-only hydropathy scales miss. We demonstrate this across three protein systems: on an -synuclein amyloid filament, strongly dewetting interfaces align with unassigned peptide densities, revealing hidden binding sites; in calmodulin, hydrophobicity redistributes upon Ca2+ binding; and for Protein G, time-resolved hydrophobicity changes track the folding trajectory. Together, these models make Fdewet a computationally inexpensive descriptor for proteins, membranes, and other surfaces, enabling rapid scoring for materials design and a time-resolved view of dynamic hydrophobic-mediated processes such as protein folding. Significance StatementHydrophobicity, the tendency of surfaces to expel water, drives how proteins fold and how molecules recognize one another. For decades, it has widely been treated as a fixed property of an amino acid or chemical group, but water actually responds to the collective shape and chemistry of a surface, such as that presented by a protein, not to its components in isolation. Measuring this collective response from molecular simulation is rigorous but prohibitively slow. We show that it can instead be inferred from a compact set of features describing water structure and interactions near a surface, and we use this insight to build models that predict hydrophobicity rapidly and at residue resolution, enabling practical, physically grounded design of hydrophobic-mediated interactions.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Universal protein misfolding intermediates can bypass the proteostasis network and remain soluble and less-functional 96%
- Deep-Learning Structure Elucidation from Single-Mutant Deep Mutational Scanning 95%
- Physics-based inverse design of cholesterolattracting transmembrane helices reveals aparadoxical role of hydrophobic length 95%
Similar papers in this journal
- Determinants of sugar-induced influx in the mammalian fructose transporter GLUT5 94%
- Differential ion dehydration energetics explains selectivity in the non-canonical lysosomal K+ channel TMEM175 94%
- Uncovering Protein Ensembles: Automated Multiconformer Model Building for X-ray Crystallography and Cryo-EM 94%
Similar papers in this journal
- COLLAPSE: A representation learning framework for identification and characterization of protein structural sites 96%
- ConforFold Recovers Alternative Protein Conformations Beyond MSA Subsampling 95%
- Neural Network-Derived Potts Models for Structure-Based Protein Design using Backbone Atomic Coordinates and Tertiary Motifs 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.