Back

InGene: Finding influential genes from embeddings of nonlinear dimension reduction techniques

Goswami, C.; Sengupta, D.

2023-06-21 bioinformatics
10.1101/2023.06.19.545592 bioRxiv
Show abstract

We introduce InGene, the first of its kind, fast and scalable non-linear, unsupervised method for analyzing single-cell RNA sequencing data (scRNA-seq). While non-linear dimensionality reduction techniques such as t-SNE and UMAP are effective at visualizing cellular sub-populations in low-dimensional space, they do not identify the specific genes that influence the transformation. InGene addresses this issue by assigning an importance score to each expressed gene based on its contribution to the construction of the low-dimensional map. InGene can provide insight into the cellular heterogeneity of scRNA-seq data and accurately identify genes associated with cell-type populations or diseases, as demonstrated in our analysis of scRNA-seq datasets.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.