BacSC: A general workflow for bacterial single-cell RNA sequencing data analysis
Ostner, J.; Kirk, T.; Olayo-Alarcon, R.; Thöming, J. G.; Rosenthal, A. Z.; Häussler, S.; Müller, C. L.
Show abstract
Bacterial single-cell RNA sequencing has the potential to elucidate within-population heterogeneity of prokaryotes, as well as their interaction with host systems. Despite conceptual similarities, the statistical properties of bacterial single-cell datasets are highly dependent on the protocol, making proper processing essential to tap their full potential. We present BacSC, a fully data-driven computational pipeline that processes bacterial single-cell data without requiring manual intervention. BacSC performs data-adaptive quality control and variance stabilization, selects suitable parameters for dimension reduction, neighborhood embedding, and clustering, and provides false discovery rate control in differential gene expression testing. We validated BacSC on a broad selection of bacterial single-cell datasets spanning multiple protocols and species. Here, BacSC detected subpopulations in Klebsiella pneumoniae, found matching structures of Pseudomonas aeruginosa under regular and low-iron conditions, and better represented subpopulation dynamics of Bacillus subtilis. BacSC thus simplifies statistical processing of bacterial single-cell data and reduces the danger of incorrect processing.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- treeclimbR pinpoints the data-dependent resolution of hierarchical hypotheses 96%
- scCDC: a computational method for gene-specific contamination detection and correction in single-cell and single-nucleus RNA-seq data 95%
- Biology-inspired data-driven quality control for scientific discovery in single-cell transcriptomics 95%
Similar papers in this journal
- Deciphering the Biosynthetic Potential of Microbial Genomes Using a BGC Language Processing Neural Network Model 95%
- Inferring cell diversity in single cell data using consortium-scale epigenetic data as a biological anchor for cell identity 95%
- CelLink: integrating single-cell multi-omics data with weak feature linkage and imbalanced cell populations 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.