Back

CSOA: A Novel Single-Cell Gene Set Enrichment Analysis Method with Comprehensive Benchmarking

Stoica, A.-F.; Yao, K.; Wang, J.; Xu, X.

2026-07-21 bioinformatics
10.64898/2026.07.16.738892 bioRxiv
Show abstract

Single-cell gene set enrichment analysis is widely used to evaluate the activity of gene sets in individual cells, as measured by single-cell sequencing technologies. However, existing methods often generate ambiguous scores that cannot reliably distinguish cells enriched for a biological signal from background cells. To address this limitation, we developed Cell Set Overlap Analysis (CSOA), a novel method for gene set enrichment analysis that leverages gene pair relationships by quantifying pairwise overlaps between high-expression cell sets constructed for each signature gene. We benchmarked CSOA against sixteen established methods representing five methodological classes: direct scoring, rank-based scoring, model-based scoring, matrix decomposition, and overrepresentation analysis. Our evaluation framework introduces novel metrics tailored for the gene set scoring problem, such as score coverage and silhouette rank alignment. They are used alongside traditional metrics for binary classification, such as the Matthews correlation coefficient and area under the receiver operating characteristic (AUROC). CSOA showed superior accurate annotation of cell types and specific biological processes compared with competing approaches. This advantage was particularly pronounced in the class boundary determination benchmark, where it ranked the first in all evaluated datasets. CSOA also outperformed most of the compared methods in computational efficiency. Notably, CSOAs combination of outstanding performance in the score coverage metric and solid overall performance positions it as a uniquely well-suited method for distinguishing cells enriched for specific biological signals.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.