Back

Statistical Causal Discovery in Developing and Refining Adverse Outcome Pathway (AOP)

Hiki, K.; Pham, T.; Yamamoto, M.; Hayashi, T. I.; Shimizu, S.

2025-09-04 biochemistry
10.1101/2025.08.31.672289 bioRxiv
Show abstract

Statistical causal discovery (SCD) has the potential to advance the development and evaluation of Adverse Outcome Pathways (AOPs) by inferring causal relationships directly from data. However, ecotoxicology data often has challenges for SCD applications, such as missing data and violation of SCD algorithm assumptions. As a proof-of-concept, we applied a linear non-Gaussian acyclic model (LiNGAM), a representative SCD method, to three types of ecotoxicology datasets: (1) bivariate dose-response relationships, (2) bivariate response- response relationships, and (3) a multivariate dataset with a known causal structure. Missing data were addressed through multiple imputation followed by causal estimation using DirectLiNGAM, a direct method for estimating LiNGAM. DirectLiNGAM identified correct causal directions with high statistical reliabilities in three of four bivariate dose-response cases, even when assumptions such as linearity and non-Gaussianity were partially violated. In contrast, response-response cases did not yield a single dominant direction, likely due to the limited number of replicates. In the multivariate case, the inferred graphs closely resembled the expert-curated causal graph, achieving high recall (0.50-0.75), despite relatively low precision (0.31-0.40). These results demonstrate the utility of SCD, combined with multiple imputation, in identifying relevant key events, revealing missing links, and refining existing AOP and quantitative AOP (qAOP) models, under realistic ecotoxicological constraints. SynopsisStatistical causal discovery can advance the development of adverse outcome pathways in a data-driven manner, enabling efficient chemical risk assessment. TOC Art O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=165 SRC="FIGDIR/small/672289v1_ufig1.gif" ALT="Figure 1"> View larger version (54K): org.highwire.dtl.DTLVardef@117bd64org.highwire.dtl.DTLVardef@19318cdorg.highwire.dtl.DTLVardef@414aecorg.highwire.dtl.DTLVardef@9e068e_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.