Back

Superior batch alignment and hyper-dimensional cytometry representations allow ultra-sensitive classification of disease phenotypes

Mashford, B. S.; Hewitt, T.; May, M.; Zhuang, Z.; Jain, A.; Diamand, K. E.; Li, F.-J.; Kwong, K.; Read, S. H.; Davies, A. R.; Hammill, D.; Andrews, T. D.

2025-07-31 systems biology
10.1101/2025.07.28.666458 bioRxiv
Show abstract

Analysis of cytometry data predominantly relies on clustering and dimensionality reduction approaches for computational tractability. This is particularly relevant for modern spectral flow cytometers, which can simultaneously measure an increasingly large number of antibody marker channels. While dimensionality reduction provides for more efficient data processing, this comes at the expense of data loss that may miss subtle patterns among rare cell types that may be critical for disease detection. Maintaining analysis at full dimensions presents opportunities to preserve resolution and provides complete downstream data interpretability. However, a significant obstacle to analysis of cytometry data without dimensionality reduction is the significant batch effects observed in this data between experiment days, operators and equipment types. Here we show a new strategy to denoise batch variation from both flow- and mass-cytometry datasets using an autoencoder neural network architecture. We generated a benchmark flow cytometry dataset in mice to compare this approach to current toolsets and find our approach shows superior preservation of biological signals, whilst also performing batch correction equal to current best methodology. Our batch alignment approach works to such an extent that it becomes practically possible to project batch aligned data into multidimensional space to generate a novel representation of cellular phenotype for downstream model building. This hyperdimensional approach maintains original data resolution without requiring dimensionality reduction, and thus any resultant cell populations that differentiate phenotypes remain fully interpretable. We show with two large clinical datasets that our batch-alignment approach coupled with the multi-dimensional representation successfully detects meaningful patterns in cases where the original analysis methods struggled. This new framework removes some of the inherent technical limitations encountered in the integration of large, multi-batch cytometry datasets and provides a framework for machine learning model building from this modality. An implementation of this framework and an associated web application accompanies this manuscript at http://voxelcoder.cloud. Significance StatementThis work addresses a critical bottleneck in cytometry analysis by introducing a neural network-based approach that effectively removes technical batch variation while preserving biological signals. The novel hyperdimensional representation maintains full data resolution without dimensionality reduction, enabling more sensitive detection of disease-associated cellular signatures than current methods. Importantly, this framework enables reliable batch normalization of fresh samples processed at different times and locations, overcoming the practical constraints of clinical sample collection where simultaneous processing is often impossible. The enhanced sensitivity for identifying pathogenic cellular patterns has immediate implications for biomarker discovery and personalized medicine applications.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Cell Reports Methods
165 papers in training set
Top 0.1%
21.8%
2
Cytometry Part A
33 papers in training set
Top 0.1%
8.9%
3
Nature Communications
5641 papers in training set
Top 21%
7.8%
4
Communications Biology
993 papers in training set
Top 0.5%
7.8%
5
npj Systems Biology and Applications
125 papers in training set
Top 0.3%
5.5%
50% of probability mass above
6
Scientific Reports
3612 papers in training set
Top 19%
5.1%
7
Patterns
78 papers in training set
Top 0.3%
4.3%
8
Briefings in Bioinformatics
354 papers in training set
Top 2%
4.0%
9
Bioinformatics
1204 papers in training set
Top 5%
3.2%
10
eLife
5828 papers in training set
Top 42%
2.4%
11
iScience
1154 papers in training set
Top 11%
2.4%
12
npj Precision Oncology
53 papers in training set
Top 0.8%
1.7%
13
BMC Bioinformatics
457 papers in training set
Top 4%
1.7%
14
Genome Biology
637 papers in training set
Top 6%
1.7%
15
Blood Advances
62 papers in training set
Top 0.8%
1.5%
16
Frontiers in Immunology
638 papers in training set
Top 8%
1.1%
17
Molecular Systems Biology
162 papers in training set
Top 2%
1.1%
18
Molecular Omics
23 papers in training set
Top 0.3%
1.1%
19
PLOS Computational Biology
1863 papers in training set
Top 17%
1.1%
20
Analytical Chemistry
218 papers in training set
Top 2%
0.9%
21
The Journal of Immunology
166 papers in training set
Top 2%
0.8%
22
Wellcome Open Research
67 papers in training set
Top 2%
0.6%
23
Bioinformatics Advances
203 papers in training set
Top 5%
0.6%
24
Nature Machine Intelligence
70 papers in training set
Top 3%
0.6%
25
Science Advances
1243 papers in training set
Top 33%
0.6%
26
PNAS Nexus
159 papers in training set
Top 4%
0.6%
27
NAR Genomics and Bioinformatics
242 papers in training set
Top 5%
0.6%