Back

CysNet: Theorem constrained inference of cysteine redox proteoform states from bottom-up mass spectrometry data

Cobley, J. N.; Jiang, H.; Platani, M.; Lamond, A. I.

2026-07-07 biochemistry
10.64898/2026.07.06.736853 bioRxiv
Show abstract

Here, we present CysNet, a theorem-constrained method designed to infer cysteine redox proteoforms (oxiforms) from bottom-up, mass spectrometry (MS) based proteomic data. This overcomes limitations with previous MS redox proteomic approaches, which can quantify residue-resolved cysteine redox states, but leave distinct oxiforms unresolved. CysNet treats each residue-resolved oxidation value as a binary redox-coordinate marginal, enabling theorem-constrained inference of the oxiforms that are necessary, impossible or bounded within the compatible protein-group ensemble. This collapses the vast theoretically possible set of oxiform states to a finite set of allowed values by extracting existence and exclusion constraints from the data, despite the incomplete proteome coverage typical for bottom-up MS datasets. Using CysNet to analyse human induced pluripotent stem cell lines (6,300 cysteine-containing protein groups, 22% cysteine coverage), resolved 519 exact oxiforms, inferring 7,000 oxiforms per line. Quantitatively, CysNet bounded the oxiform content to approx. 15% of the measured cysteine proteome. These data define the deepest oxiform survey recorded. CysNet revealed a latent structural layer in redox variation between the cell lines, distinguishing changes in oxiform identity (composition) from changes in oxiform weighting (intensity). Hence, CysNet moves bottom-up redox proteomics beyond isolated site-level cataloguing by reconstructing copy-number-weighted oxiform maps, providing a scalable route to deep oxiform information from peptide-level data.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.