Back

Multi-model reinforcement learning with online retrospective change-point detection

Chartouny, A.; Khamassi, M.; Girard, B.

2025-05-16 animal behavior and cognition
10.1101/2025.05.13.653727 bioRxiv
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWHumans continuously adapt to uncertain and changing situations. However, most reinforcement learning models of human behavior struggle to explain this capability. We propose a novel reinforcement learning agent for uncertain and volatile Markov decision processes, which we call Multi-Model with Retrospective Change Point Detection (MMRCPD). MMRCPD relies on two novel ideas: arbitrating between local models rather than contexts of the environment and retrospectively detecting change points. Arbitrating between local models limits memory costs and enables faster adaptation to new contexts which sub-parts have been experienced before. Retrospective change point detection mimics the capacity of humans to infer the latent cause of a change after it happened and maintain precise models of the environment. MMRCPD can detect local changes online, create new models, retrospectively update its models based on when it estimates that the change happened, reuse past models, merge models if they become similar, and forget unused models. This novel multi-model agent outperforms single-model and context-level change-detection methods in uncertain and locally changing environments. These results yield new insights and predictions concerning optimal decision-making in changing and uncertain environments, which could in turn be tested in behavioral experiments.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.