High-Dimensional Sensitivity Analysis for Genomic Studies: An Adversarial Framework for Learning Worst-Case Latent Confounders
Lin, Y.; Lin, K.
Show abstract
High-dimensional genomics studies are frequently confounded by unmeasured biological processes that obscure disease-specific signals. While existing workflows can estimate these latent confounders, they fail to quantify how robust a discovery is to varying levels of hypothetical confounding. We introduce sensGAN, a deep-learning adversarial framework that systematically explores the confounding spectrum by learning "worst-case" latent variables that nullify the most gene associations under novel predictive-gain constraints. By identifying the minimum confounding strength required to explain away an observed effect, our method shifts the paradigm toward a formal, quantitative sensitivity analysis. In diverse simulations, sensGAN accurately recovers latent structures and outperforms existing methods in identifying confounder-sensitive genes. Applied to human Alzheimers disease microglia, our framework prioritizes robust disease pathways while successfully isolating signals driven by unmeasured co-occurring neurodegenerative pathologies. Our method is publicly available, deposited at the GitHub repository yifanlinz/AD sensitivity ICML.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Discovering Root Causal Genes with High Throughput Perturbations 96%
- Gated recurrence enables simple and accurate sequence prediction in stochastic, changing, and structured environments 94%
- The Recurrent Temporal Restricted Boltzmann Machine Captures Neural Assembly Dynamics in Whole-brain Activity 94%
Similar papers in this journal
- Deep generative model embedding of single-cell RNA-Seq profiles on hyperspheres and hyperbolic spaces 94%
- mcRigor: a statistical method to enhance the rigor of metacell partitioning in single-cell data analysis 94%
- Learning interpretable cellular and gene signature embeddings from single-cell transcriptomic data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.