Back

PalmaClust: A graph-fusion framework leveraging the Palma ratio for robust ultra-rare cell type detection in scRNA-seq data

Niu, X.; Wang, J.; Wan, S.

2026-03-18 bioinformatics
10.64898/2026.03.16.712161 bioRxiv
Show abstract

MotivationSingle-cell RNA sequencing (scRNA-seq) is routinely used to build atlases of tissues, resolve developmental trajectories, and characterize disease microenvironments. Yet many biologically and clinically meaningful populations--including transient progenitors, therapy-resistant tumor subclones, and antigen-specific lymphocytes--occur at very low frequencies (<1%) and are easily missed by standard clustering pipelines. Existing approaches often require extensive manual curation, rely on known marker genes, or trade sensitivity for unacceptable false positive rates due to the insensitivity of metrics like the Gini index to heavy-tailed distributions. A scalable, statistically grounded method is needed to sensitively detect rare populations while providing calibrated confidence and interpretable molecular signatures. ResultsWe present PalmaClust, a graph-fusion clustering framework that repurposes Palma ratio--a tail-sensitive inequality metric in sociology--to identify marker genes driven by extreme sparsity. PalmaClust constructs and fuses multiple K-Nearest Neighbor (KNN) graphs derived from complementary gene-selection statistics including the Palma ratio, Gini index, and Fano factor. It employs a local refinement strategy that re-prioritizes Palma-ranked genes within parent clusters. Benchmarking across diverse public scRNA-seq datasets confirms that PalmaClust consistently outperforms state-of-the-art baselines, improving rare-class F1 scores by at least 20% (absolute) while maintaining high global clustering stability. Further studies demonstrate that the Palma ratio-derived graph layer is essential for capturing ultra-rare signatures that other views miss. Availabilityhttps://github.com/wan-mlab/PalmaClust.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.