SoftHybrid: A Hybrid Imputation Algorithm Optimised for Single-Cell Proteomics Data
Shi, Y.; Davis, S.; Charles, P. D.; Taylor, S.; Dombi, E.; Berridge, G.; Ebner, D.; Fischer, R.
Show abstract
Missing values (MVs) remain a significant barrier to reliable proteomics analysis, particularly in single-cell proteomics, where small amounts of starting material and limits in detection drive Missing-Not-At-Random (MNAR) sparsity. Commonly used bulk proteomic imputation approaches typically address Missing-At-Random (MAR) MVs and improve replicate consistency at the expense of sensitivity for biological variation, whereas MNAR-specific strategies preserve group differences but compromise cross-replicate reproducibility. Existing imputation methods are commonly applied to bulk data and do not offer a generalised off-the-shelf implementation that robustly addresses the significant sparsity observed in single-cell studies. Here, we introduce SoftHybrid, a continuous weighting framework that automatically balances MAR- and MNAR-oriented imputation across a dataset-derived model between missing rate and protein abundance model. Benchmarking across known ground truth samples (three-species mix) and real single-cell proteomics data showed that SoftHybrid outperforms existing methods at low inputs while meeting or exceeding previous state-of-the-art performance at the mini-bulk level. Preservation of proteomic patterns and enhanced replicate consistency improve testing significance and boost recovery of biological signals. SoftHybrid is implemented as an R package and freely available on GitHub.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- IceR improves proteome coverage and data completeness in global and single-cell proteomics 98%
- Imputation of label-free quantitative mass spectrometry-based proteomics data using self-supervised deep learning 97%
- Systematic detection of functional proteoform groups from bottom-up proteomic datasets 97%
Similar papers in this journal
Similar papers in this journal
- To fly, or not to fly, that is the question: A deep learning model for peptide detectability prediction in mass spectrometry 97%
- Real-time spectral library matching for sample multiplexed quantitative proteomics. 96%
- Automated Enrichment of Phosphotyrosine Peptides for High-Throughput Proteomics 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.