Improving SCVI for low-count cells through self-supervised augmentation
Svensson, V.
Show abstract
When analyzing single-cell RNA sequencing data with SCVI, low-UMI cells typically need to be filtered, because their learned representations lack meaningful biological signal. We show that this is caused by a specific mechanism: as UMI depth decreases, the SCVI encoder maps cells towards a learned bias point, collapsing their representations regardless of cell identity. This phenomenon is distinct from classical posterior collapse driven by KL regularization. By augmenting training with binomial thinning and adding a cross-correlation loss between original and thinned cell embeddings, the encoder learns representations that preserve cell type identity, experimental condition differences, and sample-level variation at lower UMI depths, without sacrificing reconstruction quality. These modifications extend the range of usable cells, enabling analysis of cells that would typically be discarded due to low molecule counts.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Binomial models uncover biological variation during feature selection of droplet-based single-cell RNA sequencing 96%
- Optimal tuning of weighted kNN- and diffusion-based methods for denoising single cell genomics data 95%
- Enabling interpretable machine learning for biological data with reliability scores 95%
Similar papers in this journal
- scAnnotate: an automated cell type annotation tool for single-cell RNA-sequencing data 95%
- More accurate estimation of cell composition in bulk expression through robust integration of single-cell information 94%
- SCOT+: A Comprehensive Software Suite for Single-Cell alignment Using Optimal Transport 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.