Towards Superhuman Imitation Learning for Sequential Head-and-Neck Cancer Treatment Decisions
Zhang, X.; Marai, E.; Canahuat, G.; Attia, S.; Mohamed, A.; Naser, M.; Fuller, C.; Corna, F.
Show abstract
We propose a simulator-driven imitation learning framework for sequential decision making in head and neck cancer (HNC) treatment. Our method, Superhuman Policy Gradient Optimization (SPGO), integrates inverse reinforcement learning principles with policy gradient updates to derive three-stage treatment policies directly from recorded physician decisions. It leverages a pre-trained clinical simulator--combining a variational autoencoder and gradient boosting models--to generate complete, temporally consistent patient trajectories, enabling safe and reproducible training. Unlike conventional behavior cloning, SPGO optimizes a sub-dominance loss that explicitly rewards surpassing the expert across multiple clinical outcomes, including relapse at year three and patient-reported toxicities at multiple follow-up times. We systematically compare six subdominance configurations (absolute vs. relative, sum vs. max aggregation, per-feature vs. max-only updates) to assess how loss design affects convergence and treatment quality. Our best configuration--relative differences with sum aggregation and per-feature updates--achieves over 70% superhuman dominance across clinically relevant features on held-out patients. The learned policies reproduce expert decisions on acute measures while significantly reducing predicted late toxicities and relapse risk, demonstrating generalization beyond the training distribution. CCS Concepts* Applied computing [->] Health informatics; * Computing methodologies [->] Reinforcement learning; Learning from demonstrations.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Reinforcement learning derived chemotherapeutic schedules for robust patient-specific therapy 94%
- Generalising uncertainty improves accuracy and safety of deep learning analytics applied to oncology 94%
- Mitigating Machine Learning Bias Between High Income and Low-Middle Income Countries for Enhanced Model Fairness and Generalizability 93%
Similar papers in this journal
- Data-driven Discovery of Mathematical and Physical Relations in Oncology Data using Human-understandable Machine Learning 93%
- An Explainable Multi-Modal Neural Network Architecture for Predicting Epilepsy Comorbidities Based on Administrative Claims Data 91%
- Deep Active Inference and Scene Construction 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.