Topology Matters: The Trade-off Between Wasserstein Critics and Discriminators in Single-Cell Data Integration
Reid, K.; Stein-O'Brien, G.; Guven, E.
Show abstract
MotivationHigh-throughput gene expression measurements are biased by technical and biological confounding variables, which obscure the true biological signals of interest. Adversarial autoencoders are a popular solution that corrects for these confounding effects, often relying on discriminator networks that approximate the Jensen-Shannon divergence. However, previous research has established that the Jensen-Shannon divergence suffers from vanishing gradients when distributions do not overlap, leading to poor results. We identify that this is a critical vulnerability in gene expression deconfounding, as samples from distinct experimental batches often occupy disjoint regions of the high-dimensional space. In contrast, the Wasserstein distance remains a valid metric with informative gradients even for disjoint distributions. While both approaches appear in the literature, no study has rigorously isolated the adversarial objective to systematically evaluate its impact on batch alignment, biological conservation, and scalability across varying dataset complexities. ResultsWe introduce a multi-class reference-based Wasserstein critic to systematically benchmark adversarial objectives. In fundamental binary integration tasks, we confirm that the Wasserstein critic yields superior mixing in scenarios with disjoint support. However, extensive reference sensitivity analysis reveals that this performance relies on a topologically dense reference batch. Consequently, we define the topological boundaries for adversarial alignment: the Wasserstein critic is superior for reference-mapping tasks involving dense anchors and disjoint query batches, whereas the standard discriminator approximates a global centroid more effectively in fragmented scenarios. Availability and Implementation: Source code is available at https://github.com/kreid415/wasserstein-critic-deconfounding. Data are available at https://figshare.com/articles/dataset/Benchmarking_atlas-level_data_integration_in_single-cell_genomics_-_integration_task_datasets_Immune_and_pancreas_/12420968/1. Contactkreid20@jh.edu. Supplementary informationNo supplementary data submitted.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Evaluation of out-of-distribution detection methods for data shifts in single-cell transcriptomics 95%
- scValue: value-based subsampling of large-scale single-cell transcriptomic data for machine and deep learning tasks 95%
- Sincast: a computational framework to predict cell identities in single cell transcriptomes using bulk atlases as references 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.