Back

Learning millisecond protein dynamics from what is missing in NMR spectra

Wayment-Steele, H. K.; El Nesr, G.; Hettiarachchi, R.; Kariyawasam, H.; Ovchinnikov, S.; Kern, D.

2025-03-19 biophysics
10.1101/2025.03.19.642801 bioRxiv
Show abstract

Many proteins biological functions rely on interconversions between multiple conformations occurring at micro- to millisecond ({micro}s-ms) timescales. A lack of standardized, large-scale experimental data has hindered obtaining a more predictive understanding of these motions. After curating >100 Nuclear Magnetic Resonance (NMR) relaxation datasets, we realized an observable for {micro}s-ms dynamics might be hiding in plain sight. Millisecond dynamics can cause NMR signals to broaden beyond detection, leaving some residues not assigned in the chemical shift datasets of [~]10,000 proteins deposited in the Biological Magnetic Resonance Data Bank (BMRB)1. We made the bold assumption that residues missing assignments are exchange-broadened due to {micro}s-ms motions and trained various deep learning models to predict missing assignments. Strikingly, these models also predict exchange measured via NMR relaxation experiments, indicative of {micro}s-ms dynamics. The best of these models, which we named Dyna-1, leverages an intermediate layer of the multimodal language model ESM-32. Notably, dynamics directly linked to biological function -- including enzyme catalysis and ligand binding -- are particularly well predicted by Dyna-1, which parallels our findings that residues experiencing {micro}s-ms exchange are more conserved. We anticipate the datasets and models presented here will be transformative in unlocking the common language of dynamics and function.

Published in Nature (predicted rank #3) · training set

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.