DPHL v2: An updated and comprehensive DIA pan-human assay library for quantifying more than 14,000 proteins.
Xue, Z.; Zhu, T.; Zhang, F.; Zhang, C.; Xiang, N.; Qian, L.; Yi, X.; Sun, Y.; Liu, W.; Cai, X.; Wang, L.; Dai, X.; Yue, L.; Li, L.; Pham, T. V.; Piersma, S. R.; Xiao, Q.; Luo, M.; Lu, C.; Zhu, J.; Zhao, Y.; Wang, G.; Xiao, J.; Liu, T.; Liu, Z.; He, Y.; Wu, Q.; Gong, T.; Zhu, J.; Zheng, Z.; Ye, J.; Li, Y.; Jimenez, C. R.; A, J.; Guo, T.
Show abstract
A comprehensive pan-human spectral library is critical for biomarker discovery using mass spectrometry (MS)-based proteomics. DPHL v1, a previous pan-human library built from 1096 data-dependent acquisition (DDA) MS data of 16 human tissue types, allows quantifying 10,943 proteins. However, a major limitation of DPHL v1 is the lack of semi-tryptic peptides and protein isoforms, which are abundant in clinical specimens. Here, we generated DPHL v2 from 1608 DDA-MS data acquired using Orbitrap mass spectrometers. The data included 586 DDA-MS newly acquired from 17 tissue types, while 1022 files were derived from DPHL v1. DPHL v2 thus comprises data from 24 sample types, including several cancer types (lung, breast, kidney, and prostate cancer, among others). We generated four variants of DPHL v2 to include semi-tryptic peptides and protein isoforms. DPHL v2 was then applied to a publicly available colorectal cancer dataset with 286 DIA-MS files. The numbers of identified and significantly dysregulated proteins increased by at least 21.7% and 14.2%, respectively, compared with DPHL v1. Our findings show that the increased human proteome coverage of DPHL v2 provides larger pools of potential protein biomarkers.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Mapping Protein-Protein Interactions Using Data-Dependent Acquisition Without Dynamic Exclusion 96%
- Integral-Omics: serial extraction and profiling of metabolome, lipidome, genome, transcriptome, whole proteome and phosphoproteome using biopsy tissue 96%
- TopLib: Building and searching top-down mass spectral libraries for proteoform identification 95%
Similar papers in this journal
- Deep Learning-based Pseudo-Mass Spectrometry Imaging Analysis for Precision Medicine 95%
- SingleFrag: A deep learning tool for MS/MS fragment and spectral prediction and metabolite annotation 94%
- ChemEmbed: A deep learning framework for metabolite identification using enhanced MS/MS data and multidimensional molecular embeddings 94%
Similar papers in this journal
- DeepRTAlign: toward accurate retention time alignment for large cohort mass spectrometry data analysis 96%
- Ultradeep N-glycoproteome Atlas of Mouse Reveals Spatiotemporal Signatures of Brain Aging and Neurodegenerative Diseases 95%
- Lipidomic profiling of human serum enables detection of pancreatic cancer 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.