Multi-model reinforcement learning with online retrospective change-point detection
Chartouny, A.; Khamassi, M.; Girard, B.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWHumans continuously adapt to uncertain and changing situations. However, most reinforcement learning models of human behavior struggle to explain this capability. We propose a novel reinforcement learning agent for uncertain and volatile Markov decision processes, which we call Multi-Model with Retrospective Change Point Detection (MMRCPD). MMRCPD relies on two novel ideas: arbitrating between local models rather than contexts of the environment and retrospectively detecting change points. Arbitrating between local models limits memory costs and enables faster adaptation to new contexts which sub-parts have been experienced before. Retrospective change point detection mimics the capacity of humans to infer the latent cause of a change after it happened and maintain precise models of the environment. MMRCPD can detect local changes online, create new models, retrospectively update its models based on when it estimates that the change happened, reuse past models, merge models if they become similar, and forget unused models. This novel multi-model agent outperforms single-model and context-level change-detection methods in uncertain and locally changing environments. These results yield new insights and predictions concerning optimal decision-making in changing and uncertain environments, which could in turn be tested in behavioral experiments.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Predicting human decision making in psychological tasks with recurrent neural networks 95%
- Estimating the impact of interventions against COVID-19: from lockdown to vaccination 94%
- A Hessian-based decomposition characterizes how performance in complex motor skills depends on individual strategy and variability 94%
Similar papers in this journal
- Expectation Violations as an Effective Alternative to Complex Mentalizing in Novel Communication 95%
- Coherently Remapping Toroidal Cells But Not Grid Cells are Responsible for Path Integration in Virtual Agents 94%
- WGT: Tools and algorithms for recognizing, visualizing and generating Wheeler graphs 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.