Back

Marker gene fishing for single-cell data with complex heterogeneity

Shao, Y.; Gao, Q.; Wang, L.; Li, D.; Nixon, A.; Chan, C.; Li, Q.-J.; Xie, J.

2024-11-06 bioinformatics
10.1101/2024.11.03.621735 bioRxiv
Show abstract

In single-cell studies, cells can be characterized with multiple sources of heterogeneity such as cell type, developmental stage, cell cycle phase, activation state, and so on. In some studies, many nuisance sources of heterogeneity (SOH) are of no interest, but may confound the identification of the SOH of interest, and thus affect the accurate annotate the corresponding cell subpopulations. In this paper, we develop B-Lightning, a novel and robust method designed to identify marker genes and cell subpopulations correponding to a SOH (e.g., cell activation status), isolating it from other sources of heterogeneity (e.g., cell type, cell cycle phase). B-Lightning uses an iterative approach to enrich a small set of trustworthy marker genes to more reliable marker genes and boost the signals of the SOH of interest. Multiple numerical and experimental studies showed that B-Lightning outperforms existing methods in terms of sensitivity and robustness in identifying marker genes. Moreover, it increases the power to differentiate cell subpopulations of interest from other heterogeneous cohorts. B-Lightning successfully identified new senescence markers in ciliated cells from human idiopathic pulmonary fibrosis (IPF) lung tissues, new T cell memory and effector markers in the context of SARS-COV-2 infections, and their synchronized patterns which were previously neglected. This paper highlights B-Lightnings potential as a powerful tool for single-cell data analysis, particularly in complex data sets where sources of heterogeneity of interest are entangled with numerous nuisance factors.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.