Back

Accurate Cell Abundance Quantification using Multi-positive and Unlabeled Self-learning

Lin, Y.; Chen, X.; Lin, Y.; Xiao, X.; Huang, Z.; Yang, W.; Yu, R.; Han, J.

2024-10-15 bioinformatics
10.1101/2024.10.12.617956 bioRxiv
Show abstract

Quantifying the abundance of different cell types in pathological samples can help to uncover the correlations between cell composition and pathological conditions, offering deeper insights into the roles of different cell types in complex diseases. Conventional methods for cell abundance estimation often employ unsupervised clustering or supervised learning to identify cell types and estimate their proportions. However, these methods face challenges in accurately quantifying cell abundances, as clustering results could be unreliable and supervised methods may misclassify cell types not presented in the training data. We introduce CleverXMBD1 (denoted as Clever) for quantifying cell abundance from complex samples using a multi-class classifier trained with a confidence-based multi-positive and unlabeled loss function. Our evaluations show that Clever consistently and substantially outperforms existing methods in quantifying cell type abundance across multiple single cell datasets derived from different modalities, including CyTOF and image mass cytometry.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.