Complete NMR assignment for 275 of the most common dipeptides in intrinsically disordered proteins
Rindfleisch, T.; Taule, E. F.; Miettinen, M. S.; Underhaug, J.
Show abstract
Accurate NMR chemical shift assignments are essential for atomic-resolution characterization of proteins. Especially for intrinsically disordered proteins (IDPs) and regions (IDRs), however, the assignment remains a labor-intensive task due to spectral overlap and conformational heterogeneity. Consequently, complete side-chain assignments are rare. Here, we present a comprehensive reference dataset, comprising the complete NMR chemical shift assignments for 275 of the most prevalent dipeptides in the IDPome, covering 93% of it. The dataset contains all NMR-accessible backbone and side-chain nuclei, in total 9 408 validated data points, as well as the 1D (1H, 13C) and 2D (1H-15N HSQC, 1H-13C HSQC, TOCSY, NOESY, 1H-13C HMBC) spectra used for the assignment, making it a rich resource for the training, testing, and benchmarking of tools for data-driven protein assignment, peak picking, and synthetic spectrum generation. To facilitate such machine learning applications, all data are delivered in standardized, machine-readable formats.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Cryo2StructData: A Large Labeled Cryo-EM Density Map Dataset for AI-based Modeling of Protein Structures 90%
- An interactive mass spectrometry atlas of histone posttranslational modifications in T-cell acute leukemia 89%
- A peptide-centric quantitative proteomics dataset for the phenotypic assessment of Alzheimer's disease 88%
Similar papers in this journal
Similar papers in this journal
- Slow conformational changes in the rigid and highly stable chymotrypsin inhibitor 2 93%
- Native dynamics and allosteric responses in PTP1B probed by high-resolution HDX-MS 93%
- Human Cells for Human Proteins: Isotope Labeling in Mammalian Cells for Functional NMR Studies of Disease-Relevant Proteins 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.