Back

Supervised Deep Learning with Gene Annotation for Cell Classification

Lin, Z.; Sun, W.

2024-07-15 bioinformatics
10.1101/2024.07.15.603527 bioRxiv
Show abstract

Gene-by-gene differential expression analysis is a widely used supervised approach for interpreting single-cell RNA-sequencing (scRNA-seq) data. However, modern scRNA-seq datasets often contain large numbers of cells, which can produce numerous differentially expressed genes with exceedingly small p-values but minimal effect sizes, and thus making biological interpretation difficult. To overcome this challenge, we developed Supervised Deep learning with gene ANnotation (SDAN), a method that integrates gene-annotation information with gene-expression profiles using a graph neural network. SDAN identifies functionally coherent gene sets that best classify cells, and the resulting cell-level classification scores can be aggregated to make individual-level predictions. We evaluated SDAN and two representative existing methods in three real-data applications to identify gene sets associated with severe COVID-19, dementia, and immunotherapy response in cancer. SDAN consistently outperformed alternative approaches by achieving two key objectives simultaneously: accurate classification of outcomes and unambiguous assignment of genes to gene sets of functionally related genes.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.