Back

DAWN-SCAPE: Discovering Associations With Networks through Shared Contextual Analysis of Phenotype and Expression

Shen, M.; Tian, J.; Devlin, B.; Roeder, K.

2025-12-11 genomics
10.64898/2025.12.09.691744 bioRxiv
Show abstract

The molecular basis of phenotypes is often explored by contrasting gene expression from relevant tissue taken from individuals classified into phenotypic extremes, such as affected versus unaffected individuals. Analysis of this differential expression (DE) typically identifies many genes of interest. However, it is not clear which genes differ between extremes because they alter phenotypic liability and which show differences as a result of the extreme phenotype itself. We propose a formal model to distinguish between genes that are upstream and "cause" differential expression versus those for which differential expression is a result of the initial manifestation of phenotype. Relying on two sets of p-values, one from differential expression analysis and one from gene-specific association with phenotype (AP), and a gene coexpression or other gene-based network that serves as a bridge, our method identifies communities of genes more likely upstream or downstream of the phenotype. Our method consists of three major steps: 1) gene network construction, 2) evaluation of DE and AP signal within the network to infer hidden states, and 3) detection of gene communities. We apply our method to data that were generated to assess the biological basis of autism spectrum disorder (ASD) and Alzheimers disease (AD). Our results highlight neuronal and synaptic biology as being upstream of ASD, whereas downstream processes are all non-neuronal. For AD, our results are consistent with existing hypotheses; yet, they also lend support for a recent unifying hypothesis involving cofilin/actin biology.

Matching journals

The top 11 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.