Back

BioBatchNet: A Dual-Encoder Framework for Robust Batch Effect Correction in Imaging Mass Cytometry

Liu, H.; Zhang, S.; Mao, S.; Zhao, Q.; Zhou, Y.; Gilmore, A. P.; A. Alvarez, M.; Zhou, H.

2025-03-17 bioinformatics
10.1101/2025.03.15.643447 bioRxiv
Show abstract

MotivationImaging Mass Cytometry (IMC) is a cutting-edge technology for analysing spatially resolved protein expression at the single-cell level. However, its downstream analyses are often hindered by batch effects, which introduce systematic biases and obscure true biological variations. Existing correction methods, largely developed for scRNA-seq data, struggle to achieve precise control, leading to either over-correction by removing critical biological information, or under-correction by leaving residual batch effects. Moreover, these methods face challenges in adapting to IMC data due to differences in data characteristics. Furthermore, IMC data often feature imbalanced and overlapping cell populations, complicating clustering and downstream analysis. These challenges underscore the need for a robust and controllable batch effect correction approach tailored to IMC data. ResultsWe present BioBatchNet, a dual-encoder framework utilising adversarial training to explicitly disentangle batch-specific and biological signals. BioBatchNet enables controllable batch effect correction, effectively balancing correction with the preservation of biological variation. We evaluated BioBatchNet on three IMC datasets, where it outperformed seven benchmarking methods with robustness in both correction and biological signal conservation. Additionally, we developed a Constrained Pairwise Clustering (CPC) method, which employs constrained pairs to improve clustering performance, even in datasets with imbalanced and overlapping cell populations. To validate its generalisability, BioBatchNet was also applied to four scRNA-seq datasets, where it delivered competitive performance compared to eight typical methods. These results demonstrate BioBatchNets generalisability and robustness in correcting batch effects across diverse single-cell datasets and underscore its potential for large-scale biological analyses.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.