Back

Incorporating Surfaced-Induced Dissociation Mass Spectrometry Data into an AlphaFold-derived deep learning network improves protein structure prediction

Bolz, R. M.; Day, E. H.; Drake, Z. C.; Harvey, S. R.; Wysocki, V. H.; Lindert, S.

2026-06-29 biochemistry
10.64898/2026.06.26.734850 bioRxiv
Show abstract

Surface-Induced Dissociation native Mass Spectrometry (SID-nMS) is a tandem MS activation method that yields information on the connectivity and stoichiometry of protein complexes. While insufficient for direct structure elucidation, the data derived from SID-nMS has considerable potential to inform multimeric protein structure prediction. We hypothesized that incorporating this data into a machine-learning framework could improve multimer prediction accuracy beyond that of existing deep-learning methods. To this end, we developed SIDFold, a novel AlphaFold-based deep-learning network. SIDFold is the first AlphaFold-like network to leverage experimental data during protein complex prediction, and the first deep-learning network to utilize nMS data for structure prediction. We benchmarked SIDFold on the BETA protein set, and observed an improvement in RMSD in 138 of 227 cases including 27 targets in which the predicted structure attained near-native accuracy. We then evaluated the network on 20 proteins with experimental SID-nMS data, yielding an improved RMSD in 18 cases, with five of these cases improving to a high-accuracy complex. Finally, we tested SIDFold against a previously published SID-guided Rosetta docking method, where we saw improvement in 13 of 16 proteins. SIDFold is freely available on GitHub, with example files and commands available in the Supplementary Information.

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

1
Nature Communications
5641 papers in training set
Top 16%
11.6%
2
Nature Methods
385 papers in training set
Top 1%
7.7%
3
Journal of Proteome Research
234 papers in training set
Top 0.5%
7.1%
4
Bioinformatics
1204 papers in training set
Top 5%
4.2%
5
Analytical Chemistry
218 papers in training set
Top 0.9%
3.9%
6
Molecular & Cellular Proteomics
158 papers in training set
Top 0.6%
3.3%
7
eLife
5828 papers in training set
Top 35%
3.2%
8
Angewandte Chemie International Edition
93 papers in training set
Top 0.6%
3.1%
9
Communications Biology
993 papers in training set
Top 6%
3.1%
10
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 18%
3.1%
50% of probability mass above
11
Protein Science
246 papers in training set
Top 1%
2.7%
12
Journal of the American Chemical Society
217 papers in training set
Top 1%
2.4%
13
Journal of the American Society for Mass Spectrometry
37 papers in training set
Top 0.2%
2.3%
14
Cell Reports Methods
165 papers in training set
Top 1%
2.3%
15
Structure
193 papers in training set
Top 1%
2.3%
16
ACS Chemical Biology
167 papers in training set
Top 1%
2.1%
17
PLOS Computational Biology
1863 papers in training set
Top 13%
2.1%
18
Communications Chemistry
48 papers in training set
Top 0.4%
2.1%
19
Molecular & Cellular Proteomics
25 papers in training set
Top 0.2%
2.0%
20
Journal of Chemical Information and Modeling
238 papers in training set
Top 2%
1.9%
21
ACS Omega
105 papers in training set
Top 1%
1.9%
22
Chemical Science
73 papers in training set
Top 1.0%
1.7%
23
PLOS ONE
5266 papers in training set
Top 50%
1.6%
24
Cell Systems
201 papers in training set
Top 3%
1.5%
25
Nature Machine Intelligence
70 papers in training set
Top 2%
1.1%
26
Journal of Structural Biology
64 papers in training set
Top 0.6%
1.1%
27
JACS Au
43 papers in training set
Top 0.6%
1.1%
28
Scientific Reports
3612 papers in training set
Top 70%
1.0%
29
Nucleic Acids Research
1281 papers in training set
Top 13%
1.0%
30
ACS Central Science
71 papers in training set
Top 1%
1.0%