Biologically Informative NA Deconvolution (BIND) excavates hidden features of the proteome from missing values in large-scale datasets
GUO, W.; JIN, W.; ZHENG, J.; PAN, Y.; WANG, R.; ZHANG, J.; FENG, X.; CHEN, L.; ZHANG, L.
Show abstract
The fast-advancing mass spectrometry and related technologies have greatly extended the depth of coverage in large-scale proteomics studies, including single-cell applications. As sample numbers grow rapidly, it is often challenging to interpret the proteins with missing values that are often presented as "NA" (not available). It could be the evidence of no expression, low expression below the detection threshold, or false negative detection due to technical issues. Existing methods for missing values imputation, while generally useful, rarely consider the non-random NA values that inform biological significance. In the current study, we developed Biologically Informative NA Deconvolution (BIND) that applies an adaptive neighborhood-based modeling to deconvolve the nature of NAs as "biological" (low/no expression) or technical (experimental errors). Applying to multiple cell line datasets and human tissue extracellular vesicle datasets, BIND excavated the NAs that indicated "hallmark absence" of unique proteins. This led to improvements in protein-protein interaction analysis and the identification of novel disease biomarkers. To facilitate its public accessibility, we compiled BIND into a web server that features functional online operations and interactive visualizations. Furthermore, we demonstrated that the BIND server could deconvolve the NAs and improve the analyses of single-cell proteomics datasets. Overall, BIND delineates the biological significance of missing values rather than treating them as a burden, providing a critical perspective for understanding the complex proteome in various biological contexts. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=144 SRC="FIGDIR/small/660508v1_ufig1.gif" ALT="Figure 1"> View larger version (52K): org.highwire.dtl.DTLVardef@1a9d7c4org.highwire.dtl.DTLVardef@194b2d9org.highwire.dtl.DTLVardef@169d3bborg.highwire.dtl.DTLVardef@cba0de_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Analysis and visualization of quantitative proteomics data using FragPipe-Analyst 97%
- Machine learning on large-scale proteomics data identifies tissue- and cell type-specific proteins 96%
- To fly, or not to fly, that is the question: A deep learning model for peptide detectability prediction in mass spectrometry 96%
Similar papers in this journal
- Monitoring Functional Post-Translational Modifications Using a Data-Driven Proteome Informatic Pipeline 97%
- Parallel Analyses by Mass Spectrometry (MS) and Reverse Phase Protein Array (RPPA) Reveal Complementary Proteomic Profiles in Triple-Negative Breast Cancer (TNBC) Patient Tissues and Cell Cultures 96%
- OmixLitMiner 2: Guided Literature Mining Tools for Automated Categorization of Marker Candidates in Omics Studies 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.