Unified Probabilistic Analysis of CyTOF: A Deep Generative Approach using CytoOne
Yang, Y.; Wang, K.; Shen, Y.; Weidanz, J. A.; Xiao, G.; Wang, X.
Show abstract
Extracting meaningful biological signals from Cytometry by time-of-flight (CyTOF) data remains challenging due to heterogeneity, data characteristics, and the presence of various technical artifacts. Current analysis workflows typically rely on task-specific tools assembled into pipelines, which often make inconsistent distributional assumptions and fail to fully leverage the structure of the data. We present CytoOne, a unified probabilistic framework tailored for CyTOF data that integrates batch correction, differential analysis, and visualization within a single model. CytoOne is built upon a Bayesian hierarchical architecture inspired by Nouveau Variational Autoencoders (NVAE) and employs a novel quasi zero-inflated softplus-normal (QZIPN) likelihood to flexibly model the sparse and noisy nature of CyTOF measurements. We demonstrate via qualitative and quantitative evaluations that CytoOne effectively approximates both marginal and joint distributions of CyTOF data, removes batch-specific artifacts, enables fine-grained differential expression analysis, and facilitates interpretable embeddings for exploratory analysis.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Cofea: correlation-based feature selection for single-cell chromatin accessibility data 95%
- INTREPPPID - An Orthologue-Informed Quintuplet Network for Cross-Species Prediction of Protein-Protein Interaction 94%
- Evaluating discrepancies in dimensionality reduction for time-series single-cell RNA-sequencing data 94%
Similar papers in this journal
- eSVD-DE: Cohort-wide differential expression in single-cell RNA-seq data using exponential-family embeddings 95%
- scConsensus: combining supervised and unsupervised clustering for cell type identification in single-cell RNA sequencing data 94%
- CoSTA: Unsupervised Convolutional Neural Network Learning for Spatial Transcriptomics Analysis 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.