Back

Ontology-Guided Pathway Activity Identifies a Cell-Intrinsic Defense Response Program Associated with MEK Inhibitor Sensitivity

Ndubuisi, C. W.

2026-07-28 cancer biology
10.64898/2026.07.25.740732 bioRxiv
Show abstract

Predicting cancer drug response from gene expression requires models that expose which biological pathways drive cell-line-specific sensitivity. We introduce Gene-Ontology Pathway Attention (GOPA), whose core module computes deterministic attention weights softmax(xc {middle dot}[B] ) from expression and the column-normalized gene-term annotation matrix with no learned parameters. Applying GOPA to 542 drugs across the GDSC panel under leave-cell-line-out evaluation, we find that defense response pathways predict sensitivity to kinase inhibitors in GDSC, with the strongest and most confound-resistant signal in MEK/MAPK inhibitors. This association shows cross-assay support in PRISM (11 of 11 overlapping drugs; all p < 0.002) and retains 72% signal strength after controlling for five confounds, but was not reproduced in the gCSI panel (different response metric, smaller sample, MEK inhibitors absent), indicating the finding may be MAPK-pathway-specific and assay-dependent. On the 16-drug benchmark, XGBoost achieves the lowest RMSE (1.185); GOPA (1.226) is the strongest neural model. On the full 542-drug panel, target-encoded XGBoost matches GOPA on RMSE (1.322 vs. 1.327); GOPA achieves higher residual Pearson correlation (0.469; 95% CI [0.453, 0.485]) than target-encoded XGBoost (0.391; [0.379, 0.404]); paired difference +0.078 [0.062, 0.093]; GOPA wins on 363 of 539 drugs (67.3%); Wilcoxon p = 1.1 x 10-23). When pathway representations are evaluated with matched downstream learners, simple gene-set projections achieve equivalent prediction, indicating that GOPAs value lies in its deterministic, population-comparable pathway summaries rather than representational superiority. A controlled geometry comparison shows Poincare-ball embeddings preserve GO graph distances better than Euclidean ({rho} = 0.732 vs. 0.474) while Euclidean embeddings achieve stronger ancestor retrieval; neither geometry improves prediction ({Delta}RMSE = +0.008; p = 0.18).

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.