Back

Generative cell phenotyping with structured latent populations

Bodart, F.; De Voeght, A.; Baron, F.; Louppe, G.

2026-07-03 bioinformatics
10.64898/2026.06.30.735507 bioRxiv
Show abstract

Flow cytometry produces high-dimensional single-cell protein measurements central to immunophenotyping and clinical monitoring. Yet analysis still relies largely on manual gating, which is labour-intensive, poorly reproducible, and ill-suited to large marker panels. Existing computational approaches address classification or discovery in isolation, treating cell-type identity as a post-hoc annotation rather than as part of the generative model itself. We present MARVIN, a semi-supervised variational autoencoder that encodes the assumption that cells organise into discrete populations with continuous intra-population variability through a Gaussian mixture prior in the latent space. Because each component represents a distinct cell population, classification, discovery, and density estimation emerge as complementary views of the same representation. On public benchmarks, MARVIN matches or exceeds existing methods using as few as 10% labelled cells. Trained exclusively on healthy samples, it identifies leukaemic cells through elevated reconstruction error, providing an unsupervised anomaly detection signal. On paired stimulation data, it maintains stable population assignments while capturing condition-specific shifts in abundance and marker expression at patient-level resolution. MARVIN is open-source and designed for local deployment, adapting to institution-specific panels and instruments

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
Nature Communications
5641 papers in training set
Top 5%
22.5%
2
Nature Machine Intelligence
70 papers in training set
Top 0.2%
7.9%
3
Bioinformatics
1204 papers in training set
Top 3%
6.8%
4
Cytometry Part A
33 papers in training set
Top 0.1%
6.3%
5
Nature Methods
385 papers in training set
Top 2%
4.4%
6
Nature Biotechnology
172 papers in training set
Top 1%
4.1%
50% of probability mass above
7
PLOS Computational Biology
1863 papers in training set
Top 10%
3.2%
8
Briefings in Bioinformatics
354 papers in training set
Top 3%
3.2%
9
Genome Biology
637 papers in training set
Top 4%
2.7%
10
Cell Reports Methods
165 papers in training set
Top 1%
2.4%
11
Communications Biology
993 papers in training set
Top 9%
2.4%
12
npj Systems Biology and Applications
125 papers in training set
Top 0.8%
2.1%
13
Patterns
78 papers in training set
Top 1%
1.9%
14
Molecular Systems Biology
162 papers in training set
Top 1%
1.7%
15
Scientific Reports
3612 papers in training set
Top 53%
1.7%
16
Cell Systems
201 papers in training set
Top 3%
1.3%
17
Nucleic Acids Research
1281 papers in training set
Top 10%
1.3%
18
eLife
5828 papers in training set
Top 57%
1.1%
19
NAR Genomics and Bioinformatics
242 papers in training set
Top 3%
1.1%
20
PLOS ONE
5266 papers in training set
Top 54%
1.1%
21
New Phytologist
346 papers in training set
Top 4%
1.1%
22
Science Advances
1243 papers in training set
Top 27%
1.0%
23
Nature Genetics
286 papers in training set
Top 4%
1.0%
24
BMC Bioinformatics
457 papers in training set
Top 6%
0.9%
25
iScience
1154 papers in training set
Top 34%
0.9%
26
Genome Medicine
183 papers in training set
Top 5%
0.9%
27
BMC Methods
15 papers in training set
Top 0.2%
0.9%
28
Genome Research
468 papers in training set
Top 7%
0.6%
29
Cell Reports
1498 papers in training set
Top 29%
0.6%
30
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 44%
0.6%