Back

scKAN: Interpretable Single-cell Analysis for Cell-type-specific Gene Discovery and Drug Repurposing via Kolmogorov-Arnold Networks

He, H.; Tang, Z.; Chen, G.; Xu, F.; Hu, Y.; Feng, Y.; Wu, J.; Huang, Y.-A.; Huang, Z.-A.; Tan, K. C.

2025-02-08 bioinformatics
10.1101/2025.02.04.636408 bioRxiv
Show abstract

Single-cell analysis has revolutionized our understanding of cellular heterogeneity, yet current approaches face challenges in efficiency and interpretability. In this study, we present scKAN, a framework that leverages Kolmogorov-Arnold Networks for interpretable single-cell analysis through three key innovations: efficient knowledge transfer from large language models through a lightweight distillation strategy; systematic identification of cell-type-specific functional gene sets through KANs learned activation curves; and precise marker gene discovery enabled by KANs importance scores with potential for drug repurposing applications. The model achieves superior performance on cell-type annotation with a 6.63% improvement in macro F1 score compared to state-of-the-art methods. Furthermore, scKANs learned activation curves and importance scores provide interpretable insights into cell-type-specific gene patterns, facilitating both gene set identification and marker gene discovery. We demonstrate the practical utility of scKAN through a case study on pancreatic ductal adenocarcinoma, where it successfully identified novel therapeutic targets and potential drug candidates, including Doconexent as a promising repurposing candidate. Molecular dynamics simulations further validated the stability of the predicted drug-target complexes. Our approach offers a comprehensive framework for bridging single-cell analysis with drug discovery, accelerating the translation of single-cell insights into therapeutic applications.

Published in Genome Biology (predicted rank #12) · training set

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.