Coreset-based logistic regression for atlas-scale cell type annotation
Li, D.; Ko, Y. K.; Canzar, S.
Show abstract
Logistic regression provides a versatile and interpretable framework for analyzing single-cell and spatial omics data. It has been shown to outperform more complex models on key association tasks, such as marker gene identification, and on predictive tasks, such as cell type annotation. Logistic regression has also been used effectively for quantifying spatial proximity between cell types and has demonstrated competitive performance in predicting tissue phenotypes. As increasingly large single-cell atlases become publicly available, training logistic regression models on these datasets offers opportunities for building accurate and robust reference models but also poses major computational challenges. Here, we adapt and extend the theory of coresets for logistic regression. Specifically, we compute a previously introduced classification complexity measure using linear programming to identify omics datasets that admit coresets, i.e. random subsets of cells that preserve logistic loss. Our main theoretical finding is that this complexity measure is close to one for PCA-transformed single-cell datasets. Moreover, we prove that logistic loss is preserved under this linear transformation, suggesting a universal sampling scheme that requires only a constant number of cells per cell type to obtain accurate representations of the data. Experiments across several cell atlases demonstrate that these theoretical guarantees translate to accurate and scalable transfer of cell types and spatial niches, outperforming more complex models, including pre-trained language models and graph isomorphism networks.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- LSMMD-MA: Scaling multimodal data integration for single-cell genomics data analysis 96%
- ARTEMIS integrates autoencoders and schrodinger bridges to predict continuous dynamics of gene expression, cell population and perturbation from time-series single-cell data 95%
- SCIM: Universal Single-Cell Matching with Unpaired Feature Sets 95%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.