Transition from model-free to structure-informed decision making in dynamic environments
Yasueda, M.; Taira, M.; Akam, T.; Walton, M. E.; Doya, K.
Show abstract
Reinforcement learning theory formulates distinct decision-making strategies, including reactive model-free and deliberative model-based strategies. This study investigates how mice adjust their reinforcement learning strategies while learning decision-making in dynamic environments. Unlike previous studies that focused on behaviors after extensive training periods, we analyzed changes in learning strategies in the course of training of a two-step decision-making task with probabilistic state transition and fluctuating reward probabilities. Our statistical behavioral analysis showed that the stay-probability following common and rare transitions diverged with training, a signature of strategies that utilize knowledge of task structure. We fit various reinforcement learning strategies to behavioral data and found that structure-informed strategies became increasingly dominant in their behaviors during training. Whereas previous studies emphasized transition from goal-directed to habitual strategies after extensive training, which were often associated with model-based and model-free strategies, respectively, our results newly demonstrate a shift from model-free to structure-informed strategies in early training in mice. Author summaryReinforcement learning theory allows us to examine how we make decisions and what approaches we use to optimize rewards. Most previous research, however, has examined animal behavior only after extensive training. Here we analyzed how mice adjust their reinforcement learning strategies as they are trained in a two-step decision-making task. Initially, mice relied on reactive model-free strategies, but as training progressed, their behavior began to incorporate knowledge of task structure. While previous studies suggested transition from model-based to model-free strategies with extensive training, our study revealed the opposite in the early stage of training.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Learning the structure of the world: The adaptive nature of state-space and action representations in multi-stage decision-making 94%
- Dynamic integration of forward planning and heuristic preferences during multiple goal pursuit 93%
- Revealing the mechanisms underlying latent learning with successor representations 93%
Similar papers in this journal
- Sequential delay and probability discounting tasks in mice reveal anchoring effects partially attributable to decision noise 92%
- Adapting the novel-object recognition task to quantify how much attention rodents allocate to motivationally-salient objects 91%
- Active inference, stressors, and psychological trauma: A neuroethological model of (mal)adaptive explore-exploit dynamics in ecological context 89%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.