Back

Data-Driven Modelling of Gene Expression States in Breast Cancer and their Prediction from Routine Whole Slide Images

Dawood, M.; Eastwood, M.; Jahanifar, M.; Young, L.; Ben-Hur, A.; Branson, K.; Jones, L.; Rajpoot, N.; Minhas, F. u. A. A.

2023-04-16 pathology
10.1101/2023.04.14.536756 bioRxiv
Show abstract

Identification of gene expression state of a cancer patient from routine pathology imaging and characterization of its phenotypic effects have significant clinical and therapeutic implications. However, prediction of expression of individual genes from whole slide images (WSIs) is challenging due to co-dependent or correlated expression of multiple genes. Here, we use a purely data-driven approach to first identify groups of genes with co-dependent expression and then predict their status from (WSIs) using a bespoke graph neural network. These gene groups allow us to capture the gene expression state of a patient with a small number of binary variables that are biologically meaningful and carry histopathological insights for clinically and therapeutic use cases. Prediction of gene expression state based on these gene groups allows associating histological phenotypes (cellular composition, mitotic counts, grading, etc.) with underlying gene expression patterns and opens avenues for gaining significant biological insights from routine pathology imaging directly. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=139 SRC="FIGDIR/small/536756v1_ufig1.gif" ALT="Figure 1"> View larger version (57K): org.highwire.dtl.DTLVardef@74d0dcorg.highwire.dtl.DTLVardef@13c2708org.highwire.dtl.DTLVardef@26a6dborg.highwire.dtl.DTLVardef@194b076_HPS_FORMAT_FIGEXP M_FIG C_FIG HighlightsO_LIData-driven discovery of co-expressing gene groups in breast caner C_LIO_LIHistological imaging based prediction of gene groups via deep learning C_LIO_LIIdentification of phenotypic correlates of gene-expression in histological imaging C_LIO_LIClinical and therapeutic impact of gene groups and their visual patterns identified C_LI

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.