Double Machine Learning with Multi-Gene Shared Backgroundfor Causal Inference in Single-Cell Data: Grouping Deviation Follows a Random Walk and the Accuracy-Compute Trade-Off
Ye, W.; Jiang, X.; Shen, F.
Show abstract
In high-throughput single-cell transcriptomics (p {approx} 20,000 genes), performing double machine learning (DML) causal inference on q {approx} 5,000 target genes requires nuisance function fits that grow linearly with the number of targets (Kf cross-fitting folds, Kf = 5 or 10), far exceeding feasible computational budgets, especially with deep learning. We propose a Randomized Partition Strategy (RPS): randomly divide target genes into groups, share one background compression per group, reducing deep learning model training to q/m runs (m = group size) --- a factor of m savings. The cost of grouping is accuracy loss --- we prove that the cumulative deviation of the estimator follows a one-dimensional drift-free symmetric random walk, with diffusion variance growing linearly with group size and mean squared displacement equaling the mean squared error, so accuracy loss is predictable: m = 1 is always optimal, accuracy cost is monotonically increasing, and a small accuracy sacrifice yields m-fold compute savings. On GSE189050 SLE single-cell data (Memory B cells, n = 2120), both PCA and DL methods converge to the same conclusion, confirming the random walk mechanism is method-independent; an unexpected finding is that DL diffusion growth is only 16%, far slower than PCAs 7.4 times. This work provides a quantifiable theoretical foundation for compute strategy selection in single-cell high-dimensional causal inference.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Towards inferring causal gene regulatory networks from single cell expression measurements 95%
- TarDis: Achieving Robust and Structured Disentanglement of Multiple Covariates 94%
- Belayer: Modeling discrete and continuous spatial variation in gene expression from spatially resolved transcriptomics 94%
Similar papers in this journal
Similar papers in this journal
- LDAK-KVIK performs fast and powerful mixed-model association analysis of quantitative and binary phenotypes 93%
- Separating measurement and expression models clarifies confusion in single cell RNA-seq analysis 92%
- MultiSuSiE improves multi-ancestry fine-mapping in All of Us whole-genome sequencing data 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.