NetPolicy-RL: Network-Informed Offline Reinforcement Learning for Pharmacogenomic Drug Prioritization
Lodh, E.; Majumder, S.; Chowdhury, T.; De, M.
Show abstract
Large-scale pharmacogenomic screens provide extensive measurements of drug response across diverse cancer cell lines; however, most computational approaches emphasize point-wise sensitivity prediction or static ranking, which are poorly aligned with practical decision-making, where only a limited number of candidate drugs can be tested. We propose NetPolicy-RL, a biologically informed and decision-centric framework for pharmacogenomic drug prioritization that integrates network diffusion modeling with offline reinforcement learning. Drug selection for each cell line is formulated as an offline contextual bandit problem, enabling direct optimization of ranking quality rather than surrogate regression objectives. Mechanistic biological context is incorporated by propagating drug targets over curated interaction networks (STRING and Reactome) using random walk with restart, and combining the resulting diffusion profiles with cell-specific molecular importance derived from multi-omics data to compute network disruption scores. These biologically grounded signals are integrated with normalized drug response measurements to construct a joint state representation, which is optimized using an offline actor-critic architecture. Across held-out test splits, NetPolicy-RL consistently outperforms global ranking heuristics and learning-to-rank baselines, achieving statistically significant improvements in per-cell Normalized Discounted Cumulative Gain (NDCG@10) and substantial reductions in per-cell regret. Relative to GlobalTopK, the policy improves NDCG@10 for 88.7% of cell lines, while improvements exceed 95% compared with LambdaMART and regression-to-ranking baselines. Ablation analyses show that neither empirical response signals nor network-derived features alone are sufficient, and that their integration yields the most robust performance. Overall, this study demonstrates that combining mechanistic network biology with offline policy learning provides an effective and interpretable framework for drug prioritization in precision oncology.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Looking at the BiG picture: Incorporating bipartite graphs in drug response prediction 96%
- TUGDA: Task uncertainty guided domain adaptation for robust generalization of cancer drug response prediction from in vitro to in vivo settings 96%
- MOViDA: Multi-Omics Visible Drug Activity Prediction with a Biologically Informed Neural Network Model 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Personalized cancer treatment strategies incorporating irreversible and reversible drug resistance mechanisms 95%
- SwarmMAP: Swarm Learning for Decentralized Cell Type Annotation in Single Cell Sequencing Data 93%
- Delaying Cancer Progression by Integrating Toxicity Constraints in a Model of Adaptive Therapy 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.