Back

ANS: Adjusted Neighborhood Scoring to improve assessment of gene signatures in single-cell RNA-seq data

Ciernik, L.; Kraft, A.; Barkmann, F.; Yates, J.; Boeva, V.

2023-09-22 bioinformatics
10.1101/2023.09.20.558114 bioRxiv
Show abstract

In the field of single-cell RNA sequencing (scRNA-seq), gene signature scoring is integral for pinpointing and characterizing distinct cell populations. However, challenges arise in ensuring the robustness and comparability of scores across various gene signatures and across different batches and conditions. Here, we evaluated the stability of established methods such as Scanpy, UCell, and JASMINE in the context of scoring cells of different types and states. On eight cancer and healthy scRNA-seq datasets, we reported that none of the existing methods provide fair gene signature scores that can be used in unsupervised cell state annotation based on the highest signature values. Addressing this challenge, we introduced a new scoring method, the Adjusted Neighbourhood Scoring (ANS), that builds on the traditional Scanpy method and improves the handling of the control gene sets. We further exemplified the usability of ANS scoring in differentiating between cancer-associated fibroblasts and malignant cells undergoing epithelial-mesenchymal transition (EMT) in four cancer types, and evidenced excellent classification performance (AUCPR train: 0.95-0.99, AUCPR test: 0.91-0.99). In summary, our research introduces ANS as a robust and deterministic scoring approach that enables the comparison of diverse gene signatures and score-based annotation of cell types and states. The results of our study contribute to the development of more accurate and reliable methods for analyzing scRNA-seq data.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.