GONNECT: A Gene Ontology-guided Neural Network for Explainable Cancer Typing
Lieftinck, M.; Verlaan, T.; Reinders, M.
Show abstract
Deep Neural Networks (DNNs) are renowned for their high accuracy and versatility, which has led to their application in many fields of research, including biology. However, this accuracy often comes at the expense of interpretability, making it challenging to reason about the inner workings of most DNNs. Particularly in biological research, understanding the mechanisms behind specific outcomes is highly valuable. To elucidate the latent space of DNNs in the context of cancer biology, we introduce GONNECT: a Gene Ontology-derived Neural Network for Explainable Cancer Typing. GONNECT incorporates biological prior knowledge from the Gene Ontology (GO) directly into its network architecture, enabling interpretability through model structure. Using an autoencoder framework, we evaluate GONNECT as both encoder and decoder module and demonstrate its ability to learn which biological processes are distinctive for different cancer types. Furthermore, we show how a variant including soft links (GONNECT-SL) can expand on current knowledge by proposing new interactions between biological processes. GONNECT is flexible both in the amount of prior knowledge it incorporates and the set of input genes.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- scPRINT: pre-training on 50 million cells allows robust gene network predictions 97%
- GRouNdGAN: GRN-guided simulation of single-cell RNA-seq data using causal generative adversarial networks 97%
- stLearn: integrating spatial location, tissue morphology and gene expression to find cell types, cell-cell interactions and spatial trajectories within undissociated tissues 96%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.