Back

Improved marker detection for rare population in single-cell transcriptomics through text mining-inspired scoring approach

Chen, J.; Li, M.; Bhuva, D. D.; Davis, M. J.; Papenfuss, A. T.; Tan, C. W.

2025-11-12 bioinformatics
10.1101/2025.11.10.687759 bioRxiv
Show abstract

Accurate identification of cell types and states is a crucial step when analysing single cell RNA-seq data. Inaccurate cell type annotation leads to spurious results and biological interpretations in subsequent analyses. Cell type identification is traditionally done using established marker genes of each cell population. However, many existing methods do not perform well for imbalanced cell populations, which often occurs in the presence of rare cell types. Most existing methods tend to be biased towards the dominant cell populations or highly expressed genes, leading to inaccurate results. Here, we introduce the package smartid, which accurately identifies markers from even imbalanced data across batches by employing a modified Term Frequency-Inverse Document Frequency approach with Gaussian Mixture Model. smartid is also a gene-set scoring method which is able to distinguish the target group of interest. smartid is implemented in R and is freely available on Bioconductor at https://bioconductor.org/packages/release/bioc/html/smartid.html.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.