Back

SoftHybrid: A Hybrid Imputation Algorithm Optimised for Single-Cell Proteomics Data

Shi, Y.; Davis, S.; Charles, P. D.; Taylor, S.; Dombi, E.; Berridge, G.; Ebner, D.; Fischer, R.

2026-01-14 bioinformatics
10.64898/2026.01.13.699212 bioRxiv
Show abstract

Missing values (MVs) remain a significant barrier to reliable proteomics analysis, particularly in single-cell proteomics, where small amounts of starting material and limits in detection drive Missing-Not-At-Random (MNAR) sparsity. Commonly used bulk proteomic imputation approaches typically address Missing-At-Random (MAR) MVs and improve replicate consistency at the expense of sensitivity for biological variation, whereas MNAR-specific strategies preserve group differences but compromise cross-replicate reproducibility. Existing imputation methods are commonly applied to bulk data and do not offer a generalised off-the-shelf implementation that robustly addresses the significant sparsity observed in single-cell studies. Here, we introduce SoftHybrid, a continuous weighting framework that automatically balances MAR- and MNAR-oriented imputation across a dataset-derived model between missing rate and protein abundance model. Benchmarking across known ground truth samples (three-species mix) and real single-cell proteomics data showed that SoftHybrid outperforms existing methods at low inputs while meeting or exceeding previous state-of-the-art performance at the mini-bulk level. Preservation of proteomic patterns and enhanced replicate consistency improve testing significance and boost recovery of biological signals. SoftHybrid is implemented as an R package and freely available on GitHub.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.