PLIERv2: bigger, better and faster
Subirana-Granes, M.; Nandi, S.; Zhang, H.; Chikina, M.; Pividori, M.
Show abstract
Gene expression analysis has long been fundamental for elucidating molecular pathways and gene-disease relationships, but traditional single-gene approaches cannot capture the coordinated regulatory networks underlying complex phenotypes; although unsupervised matrix factorization methods (e.g., PCA, NMF) reveal coexpression patterns, they lack the ability to incorporate prior biological knowledge and often struggle with interpretability and technical noise correction. Semi-supervised strategies such as PLIER have improved interpretability by integrating pathway annotations during latent variable extraction, yet the original PLIER implementation is prohibitively slow and memory-intensive, making it impractical for modern large-scale resources like ARCHS4 or recount3. Here, we introduce CLAMP, which overcomes these constraints through a two-phase algorithmic design (an unsupervised "CLAMPbase" initialization followed by a "CLAMPfull" regression that incorporates priors via glmnet), rigorous internal cross-validation to tune regularization parameters for each latent variable, and efficient on-disk data handling using memory-mapped matrices from the bigstatsr package. Benchmarking on GTEx, recount2, and ARCHS4 demonstrates that CLAMP achieves 7x-41x speedups over PLIER, succeeds in modeling hundreds of thousands of samples that PLIER cannot handle, and maintains or improves biological specificity of latent variables as shown by tissue-alignment and pathway enrichment analyses. By filling the gap in scalable, biologically informed latent variable extraction, CLAMP enables comprehensive analysis of modern transcriptomic compendia and paves the way for deeper insights into gene regulatory networks and downstream applications in translational genomics.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- MetaCell: analysis of single cell RNA-seq data using k-NN graph partitions 95%
- scDesign2: a transparent simulator that generates high-fidelity single-cell gene expression count data with gene correlations captured 94%
- BANDITS: Bayesian differential splicing accounting for sample-to-sample variability and mapping uncertainty 94%
Similar papers in this journal
- eSVD-DE: Cohort-wide differential expression in single-cell RNA-seq data using exponential-family embeddings 94%
- CoSTA: Unsupervised Convolutional Neural Network Learning for Spatial Transcriptomics Analysis 94%
- Denoising Single-Cell RNA-Seq Data with a Deep Learning-Embedded Statistical Framework 93%
Similar papers in this journal
Similar papers in this journal
- CellMentor: Cell-Type Aware Dimensionality Reduction for Single-cell RNA-Sequencing Data 95%
- Constructing Ensemble Gene Functional Networks Capturing Tissue/condition-specific Co-expression from Unlabled Transcriptomic Data with TEA-GCN 94%
- MOCCASIN: A method for correcting for known and unknown confounders in RNA splicing analysis 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.