Back

GRAPEVINE: A Reinforcement Failing Framework for AI-Guided Discovery of Hidden Spatial Dynamics in Adaptive Tumor Therapy

Aif, S.; Eiche, M.; Appold, N.; Fischer, E.; Citak, T.; Kayser, J.

2025-04-10 biophysics
10.1101/2025.04.08.647768 bioRxiv
Show abstract

Artificial intelligence is revolutionizing scientific discovery in medicine, with reinforcement learning (RL) emerging as a promising tool for optimizing therapeutic strategies. Yet applying RL to complex scenarios such as therapy dynamics in solid tumors is constrained by the challenge of constructing training environments that are both computationally efficient and mechanistically interpretable. Here we introduce Reinforcement Failing, an AI-guided, human-in-the-loop discovery framework that shifts the focus from agent policy optimization to the refinement of the training environment itself. By combining multi-fidelity RL with group-relative performance evaluation across agent cohorts, Reinforcement Failing systematically reveals emergent mechanisms that first-principles models overlook. We apply this framework to adaptive therapy in solid tumors, which seeks to delay resistance-mediated treatment failure. In this setting, Reinforcement Failing uncovered a coupling between the mechanically driven collective motion of cells and spatially-heterogeneous proliferation that strongly influences therapy outcomes. Incorporating these emergent physical mechanisms into an augmented training environment improved cross-environment therapeutic performance and exposed potential pitfalls in translation. More broadly, these findings position Reinforcement Failing as a powerful artificial scientific discovery framework, capable of deciphering high-complexity processes at the interface of physics, machine learning, and medicine.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.