Back

gtexture: Haralick texture analysis for graphs and its application to biological networks

Barker-Clarke, R. J.; Scott, J. G.; Weaver, D.

2022-11-24 bioinformatics
10.1101/2022.11.21.517417 bioRxiv
Show abstract

ObjectiveThe calculation of texture features, such as those derived by Haralick et al., has been traditionally limited to 2D-imaging data. We present the novel derivation of an extension to these texture features that can be applied to graphs and networks and set out to illustrate the potential of these metrics for use in cancer informatics. ApproachWe extend the pixel-based calculation of texture and generate analogous novel metrics for graphs and networks. The graph structures in question must have ordered or continuous node weights/attributes. To demonstrate the utility of these metrics in cancer biology, we demonstrate these metrics can distinguish different fitness landscapes, gene co-expression and regulatory networks, and protein interaction networks with both simulated and publicly available experimental gene expression data. Main ResultsWe demonstrate that texture features are informative of graph structure and analyse their sensitivity to discretization parameters and node label noise. We demonstrate that graph texture varies across multiple network types including fitness landscapes and large protein interaction networks with experimental expression data. We show the ability of these texture metrics, calculated on specific protein interaction subnetworks, to classify cell line expression by lineage, generating classifiers with 82% and 89% accuracy. SignificanceGraph texture features are a novel second order graph metric that can distinguish cancer types and topologies of evolutionary landscapes. It appears that no similar metrics currently exist and thus we open up the potential derivation of more metrics for the classification and analysis of network-structured data. This may be particularly useful in the complex setting of cancer, where large graph and network structures underlie the omics data generated. Network-based data underlies drug discovery, drug response prediction and single-cell dynamics and thus these metrics provide an additional tool in tackling these problems in cancer.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.