Back

Single-cell phenotypic analysis and multiplet detection through incorporation of microscopy data into cellenONE-based single-cell proteomic data analysis

Loi, M.; Holmes, S.; Zhou, S.; Held, M.; Brownridge, P.; Coomber, J.; Neves, L. X.; Sweeney, T. R.; Emmott, E.

2025-12-02 systems biology
10.64898/2025.12.01.691577 bioRxiv
Show abstract

The reliability of single-cell proteomics (SCP) is intrinsically linked to the fidelity of cell isolation; however, the identification of co-isolated cells (doublets or multiplets) remains a persistent challenge for single-cell (sc-) omics. While single-cell transcriptomics has established probabilistic frameworks for doublet detection, these methods are ill-suited for the sparser throughput of SCP datasets. This work presents scpImaging, a novel computational pipeline that repurposes the latent microscopy data routinely generated during cellenONE-based sample preparation to provide deterministic quality control (QC) and high-content phenotypic profiling. We demonstrate that standard proteomic quality metrics for SCP (e.g., peptide count, signal intensity) paradoxically favour retaining doublets. In contrast, scpImaging applies machine-learning-based cell segmentation (Cellpose-SAM) to identify and exclude these artefacts with high precision. Furthermore, the framework integrates morphological metrics generated using the open-source CellProfiler, with proteomic abundance data, enabling joint analysis that links cell shape and texture to molecular cell state. Provided as an open-source R package, scpImaging offers a scalable, automated solution for enhancing data integrity and adding a phenotypic dimension to SCP experiments without increasing experimental cost or complexity.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.