Back

Improving SCVI for low-count cells through self-supervised augmentation

Svensson, V.

2026-02-13 bioinformatics
10.64898/2026.02.11.705441 bioRxiv
Show abstract

When analyzing single-cell RNA sequencing data with SCVI, low-UMI cells typically need to be filtered, because their learned representations lack meaningful biological signal. We show that this is caused by a specific mechanism: as UMI depth decreases, the SCVI encoder maps cells towards a learned bias point, collapsing their representations regardless of cell identity. This phenomenon is distinct from classical posterior collapse driven by KL regularization. By augmenting training with binomial thinning and adding a cross-correlation loss between original and thinned cell embeddings, the encoder learns representations that preserve cell type identity, experimental condition differences, and sample-level variation at lower UMI depths, without sacrificing reconstruction quality. These modifications extend the range of usable cells, enabling analysis of cells that would typically be discarded due to low molecule counts.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.