A Graph-Attentive GAN for Rare-Cell-Aware Single-Cell RNA-Seq Data Generation
Ganguly, R.; Aafrine, S.; Hossain, S. M. M.; Ray, S.
Show abstract
A central challenge in downstream single-cell RNA sequencing (scRNA-seq) analysis is the high-dimensional, small-sample (HDSS) regime, often compounded by class imbalance from rare cell types. These factors hinder robust feature (gene) selection and cell clustering and limit the realism of samples generated by existing simulators. We introduce GARAGE, a Graph-Attentive RAre-cell aware single-cell data GEneration that augments the generators input with a small, attention-weighted leakage of real cells in addition to prior noise. Specifically, we build a k-nearest-neighbour cell graph and use a graph attention network (GAT) to prioritize nodes that likely represent under-sampled (rare) subpopulations; these high-attention cell embeddings are injected into the generator input to steer synthesis toward biologically plausible regions of the data manifold while respecting cell-type proportions. This attention-guided leakage accelerates training, reduces mode dropping, and yields realistic synthetic cells that preserve rare-cell structure. Across real scRNA-seq benchmarks, GARAGE improves downstream feature selection and clustering compared with state-of-the-art baselines. In summary, GARAGE directly addresses HDSS and rarity in scRNA-seq by coupling graph attention with adversarial generation to produce high-fidelity synthetic cells that enhance downstream analyses.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Evaluating discrepancies in dimensionality reduction for time-series single-cell RNA-sequencing data 96%
- Synthetic observations from deep generative models and binary omics data with limited sample size 96%
- scValue: value-based subsampling of large-scale single-cell transcriptomic data for machine and deep learning tasks 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.