Accurate highly variable gene selection using RECODE in transcriptome data analysis
Imoto, Y.
Show abstract
Recent transcriptomics technologies enable gene-expression profiling at single-cell or micrometer-scale spatial resolution, but capture only a small fraction of true RNA molecules, introducing substantial technical noise driven by random sampling. These noise effects distort the earliest analytical steps, dimensionality reduction or highly variable gene (HVG) selection, and their consequences propagate into downstream analyses. The central aim of this study is to address this issue fundamentally by appropriately removing technical noise at its source. Here, I demonstrate that HVG selection based on RECODE, a de-noising method grounded in high-dimensional statistical theory, outperforms widely used approaches for both scRNA-seq and spatial transcriptomics data. RECODE-based HVG selection achieves higher accuracy and robustness, avoids missing values, improves down-stream performance, and provides the fastest runtime and best scalability among noise-reduction methods. These findings show that theory-driven noise removal is essential for recovering true biological signals and establish RECODE as a practical and reliable preprocessing strategy for single-cell analysis.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CDSeqR: fast complete deconvolution for gene expression data from bulk tissues 97%
- scConsensus: combining supervised and unsupervised clustering for cell type identification in single-cell RNA sequencing data 96%
- GEOlimma: Differential Expression Analysis and Feature Selection Using Pre-Existing Microarray Data 95%
Similar papers in this journal
- SMNN: Batch Effect Correction for Single-cell RNA-seq data via Supervised Mutual Nearest Neighbor Detection 96%
- labelSeg: segment annotation for tumor copy number alteration profiles 96%
- CosTaL: An Accurate and Scalable Graph-Based Clustering Algorithm for High-Dimensional Single-Cell Data Analysis 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.