Back

The geometry of G x E: how scaling and endogenous treatment effects shape interaction direction

Sadowski, M.; Dahl, A. W.; Zaitlen, N.; Border, R.

2025-07-18 genetics
10.1101/2025.07.15.664999 bioRxiv
Show abstract

Gene-environment interaction (G x E) studies hold promise for identifying genetic loci mediating the effects of environmental risk on disease. However, interpretation of G x E effects is often confounded by two fundamental issues: the dependence of interaction estimates on outcome scale and the presence of endogenous treatment effects, in which genetic liability influences environmental exposure. These factors can induce spurious G x E signals--even when genetic and environmental contributions are purely additive on an unobserved scale. In this work, we demonstrate that any monotone convex transformation of an outcome induces sign-consistent G x E effects: the sign of the interaction term aligns with the sign of the corresponding main genetic effect. We further show that endogenous treatment effects, modeled as threshold-based interventions, generate G x E effects with a similar directional signature. Exploiting this property, we propose a simple diagnostic: sign consistency across G x E estimates can identify artifacts driven by outcome scaling or exposure endogeneity. We validate our framework in the UK Biobank using transcriptome-wide interaction studies (TxEWAS) across multiple trait-environment pairs, observing widespread sign consistency in some settings--suggesting confounding by scaling or treatment bias. Our results provide both a theoretical foundation and a practical tool for interpreting G x E findings, enabling researchers to distinguish biologically meaningful interactions from those induced by statistical artifacts.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.