Back

Geometric-aware and interpretable deep learning for single-cell batch correction via explicit disentanglement and optimal transport

Jiang, C.; Zheng, R.; Ji, Y.; Cao, S.; Fang, Y.; Wang, Z.; Wang, R.; Liang, S.; Tao, S.

2026-02-20 bioinformatics
10.64898/2026.02.17.706490 bioRxiv
Show abstract

Single-cell RNA sequencing enables high-resolution characterization of cellular heterogeneity, yet integrating datasets from diverse sources remains challenging due to batch effects. Current methods rely on implicit feature disentanglement and and lack geometric constraints, often result in under-correction, over-correction, or compromised biological fidelity. Here, we present iDLC, an interpretable deep learning framework that performs dual-level correction through explicit feature disentanglement and optimal transport-regularized adversarial alignment. iDLC separates biological and technical components within a structured latent space, then leverages high-confidence mutual nearest neighbor pairs to guide geometrically constrained distribution alignment. Systematic evaluation across pancreatic cancer datasets with varying batch effect intensities, multi-source human immune cells, and large-scale cross-species atlases demonstrates that iDLC robustly eliminates complex batch effects while preserving fine-grained cell subtypes, continuous developmental trajectories, and rare populations. The framework scales efficiently to datasets exceeding one million cells and consistently outperforms existing methods in both batch correction and biological conservation metrics. iDLC provides a principled and reliable tool for constructing unified single-cell reference atlases across diverse experimental conditions and biological systems.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.