Learning Minimal Gene Programs for Disease-Aligned Representations
Madduri, A.; Patel, C. J.
Show abstract
Identifying small, interpretable gene sets that robustly capture disease-associated variation in singlecell transcriptomic data remains a central challenge for biological interpretation and experimental followup. In practice, commonly used differential expression and sparsity-based approaches often produce large, unstable gene lists that fail to generalize across patients due to strong donor-specific confounding. We study sparse gene selection for reconstructing donor-robust, disease-aligned cellular trajectories in real single-cell RNA-seq datasets. We introduce Sparse Linear Manifold Control (SLMC), a practical workflow that defines a disease-aligned score after removing donor-associated variation and selects minimal gene programs whose expression reconstructs this score. We focus on diagnosing the structure of the resulting reconstruction objective and evaluating selection strategies under realistic health data conditions. Across five human single-cell datasets spanning oncology and neurodegeneration, we find that the reconstruction objective exhibits strong diminishing returns, explaining why simple greedy selection methods perform well in practice. Under strict donor-heldout evaluation, greedy methods consistently outperform LASSO at small gene budgets and achieve accurate reconstruction with as few as 25 genes. Together, these results highlight how careful objective design and empirical evaluation enable robust and interpretable gene selection for disease-aligned representation learning in single-cell health data. Data and Code AvailabilityAll code used for data preprocessing, model training, evaluation, and feature importance analyses is available at: https://github.com/AdiVM/SLMC_single-cell. The singlecell data used in this study are publicly available human transcriptomic datasets generated by prior studies and accessible through the Gene Expression Omnibus (GEO). Analyses were performed using renal cell carcinoma single-cell RNA-seq data (GEO accession: GSE314072) and Alzheimers disease single-nucleus RNA-seq data from human cortex (GEO accession: GSE138852), with additional publicly available datasets used for cross-context diagnostic evaluation (GEO accessions: GSM8652069, GSE308624, and GSE227734). All datasets contain de-identified human samples and are available under standard publicuse terms via GEO. Institutional Review Board (IRB)This study analyzes de-identified, publicly available human transcriptomic data obtained from previously published studies. No new data were collected, and no identifiable private information was accessed. In accordance with institutional policy, this work was determined to constitute non-human subjects research and did not require additional IRB approval.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ARTEMIS integrates autoencoders and schrodinger bridges to predict continuous dynamics of gene expression, cell population and perturbation from time-series single-cell data 95%
- SCIM: Universal Single-Cell Matching with Unpaired Feature Sets 95%
- Identifying cancer pathway dysregulations using differential causal effects 95%
Similar papers in this journal
- An in-depth comparison of linear and non-linear joint embedding methods for bulk and single-cell multi-omics 94%
- scValue: value-based subsampling of large-scale single-cell transcriptomic data for machine and deep learning tasks 94%
- Novel multi-omics deconfounding variational autoencoders can obtain meaningful disease subtyping 94%
Similar papers in this journal
Similar papers in this journal
- Highly Accurate Cancer Phenotype Prediction with AKLIMATE, a Stacked Kernel Learner Integrating Multimodal Genomic Data and Pathway Knowledge 95%
- Optimal transport reveals dynamic gene regulatory networks via gene velocity estimation 95%
- Learning Genetic Perturbation Effects with Variational Causal Inference 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.