Back

SpinForecast: chain-free probabilistic backbone assignment of intrinsically disordered proteins from NMR chemical shifts

Eaton, J. T.; Cornish, J.; Silvey, K. M.; Chakraborty, P.; Löhr, T.; Karunanithy, G.; Heller, G. T.

2026-08-05 biophysics
10.64898/2026.07.31.740808 bioRxiv
Show abstract

Assigning peaks in NMR spectra to specific residues is an essential but time-consuming step in the study of intrinsically disordered proteins (IDPs). Conventional approaches rely on building chains of sequential connectivities between peaks, which are particularly prone to failure in disordered systems due to spectral overlap, missing peaks, and proline-rich sequences. Here we present SpinForecast, a tool that performs chain-free probabilistic backbone assignment of IDPs from chemical shifts alone, without requiring peaks to be linked into sequential chains. SpinForecast predicts residue-specific chemical shift distributions from a disorder-filtered subset of the Biomolecular Magnetic Resonance Data Bank (BMRB), incorporating nearest-neighbour residue effects and corrections for temperature and pH. Experimental chemical shifts are then assigned to residues by Bayes theorem, using residue-specific chemical shift distributions as likelihoods. We validate SpinForecast on three disordered proteins, IAPP (37 residues), NUPR1 (82 residues), and JPT2 (218 residues), achieving 100% confidence assignments for 74%, 57%, and 44% of in-distribution peaks, respectively, all with at least 99% accuracy. Where single assignments cannot be determined for these systems, the correct assignment was contained within the returned set of candidate assignments in greater than 97% of cases. SpinForecast is freely available at https://tools.bindresearch.org/bindbox/SpinForecast.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.