A hyperparameter benchmark of VAE-based methods for scRNA-seq batch integration
Kassab, M.; Maniero, L.; Beltrame, E.
Show abstract
We present the first systematic benchmark of model-architecture hyperparameters for variational autoencoder (VAE) methods for single-cell RNA-seq batch integration within scvi-tools, comparing scVI, MrVI, and LDVAE across four heterogeneous datasets under two feature regimes (all genes vs highly variable genes (HVGs)). We investigated 960 trainings (120 configurations) varying latent size and network depth/width, and evaluated with a standardized scIB metric suite covering batch removal and biological conservation (Batch ASW, PCR-batch, iLISI, graph connectivity, NMI, ARI, label ASW, isolated-label F1/ASW, cLISI, trajectory conservation), plus qualitative UMAP/t-SNE and PCA, random projection, and unintegrated baselines. Results show dataset-dependent trade-offs: scVI performs best overall via stronger batch correction; LDVAE can better preserve biological structure in some datasets; MrVI is stable and excels at batch correction in multi-protocol settings but is more resource-intensive. HVG-only training generally outperforms full-gene training for all models. Hyperparameter analysis suggests moderate-to-high latent dimensionality (>30) often gives the best balance; sensitivity to latent size tracks dataset heterogeneity (tissues, labs, chemistries, gene coverage), with larger latents improving batch mixing but sometimes reducing biological conservation. We provide model- and dataset-specific guidelines for practical defaults and tuning of VAE-based integration in single-cell studies. Reproducibility code is available on GitHub at: https://github.com/Kassab11/A-hyperparameter-benchmark-of-VAE-based-methods-for-scRNA-seq-batch-integration
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Highly Accurate Cancer Phenotype Prediction with AKLIMATE, a Stacked Kernel Learner Integrating Multimodal Genomic Data and Pathway Knowledge 95%
- Clustering and visualization of single-cell RNA-seq data using path metrics 95%
- Accessible, Reproducible, and Scalable Machine Learning for Biomedicine 95%
Similar papers in this journal
- scValue: value-based subsampling of large-scale single-cell transcriptomic data for machine and deep learning tasks 96%
- Evaluating discrepancies in dimensionality reduction for time-series single-cell RNA-sequencing data 94%
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations 94%
Similar papers in this journal
- Generating hard-to-obtain information from easy-to-obtain information: applications in drug discovery and clinical inference 96%
- Single-Cell Multi-Modal GAN (scMMGAN) reveals spatial patterns in single-cell data from triple negative breast cancer 95%
- Tokenized and Continuous Embedding Compressions of Protein Sequence and Structure 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.