Chemotherapy dose scheduling via Q-learning in a Markov tumor model
Giles, M.; Newton, P. K.
Show abstract
We describe a Q-learning approach to optimized chemotherapy dose scheduling in a stochastic finite-cell Markov process that models tumor cell natural selection dynamics. The three competing subpopulations comprising our virtual tumor are a chemo-sensitive population (S), and two chemo-resistant populations, R1 and R2, each resistant to one of two drugs, C1 and C2. The two drugs are toggled off or on which constitute the actions (selection pressure) imposed on our state-variables (S, R1, R2), measured as proportions in our finite state-space of N cancer cells (S + R1 + R2 = N). After the converged chemo-dosing policies are obtained, corresponding to a given reward structure, we focus on three important aspects of chemotherapy dose scheduling. First, we identify the most likely evolutionary paths of the tumor cell populations in response to the optimized (converged) policies. Second, we quantify the robustness in the ability to reach our target of balanced co-existence in light of incomplete information in both the initial cell-populations as well as the state-variables at each step. Third, we evaluate the efficacy of simplified policies which exploit the symmetries uncovered from an examination of the full policy. Our reward structure is designed to delay the onset of chemo-resistance in the tumor by rewarding a well-balanced mix of co-existing states, while punishing unbalanced subpopulations to avoid extinction.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Optimal Control Strategies for Mitigating Antibiotic Resistance: Integrating Virus Dynamics for Enhanced Intervention Design 96%
- The straight and narrow: a game theory model of broad- and narrow-spectrum empiric antibiotic therapy 96%
- On the design and stability of cancer adaptive therapy cycles: deterministic and stochastic models 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.