Back

Nonlinear influence of reward volatility on arbitration between multiple learning strategies reflects cost-benefit optimization

Yamada, T.; Samejima, K.

2026-06-19 animal behavior and cognition
10.64898/2026.06.15.732293 bioRxiv
Show abstract

Action selection involves two systems: a model-free reinforcement learning strategy, which relies on experience with action-outcome pairs, and a model-based reinforcement learning strategy, which enables more flexible behavior via inference using a model of the invariant environmental structure. Although environmental change requires more flexible behavior, the ability of volatility, a higher-order statistic that captures how rapidly or frequently the environment changes, to systematically modulate these strategies remains unclear. We examined the effects of reward volatility on arbitration between model-free and model-based reinforcement learning strategies using two modified two-step decision tasks. In Experiment 1, participants performed tasks with different levels of reward volatility and time pressure. In Experiment 2, we systematically manipulated reward volatility across a broader range to assess the relationship between volatility and learning strategy. Behavioral data were analyzed using model-agnostic one-trial and multitrial back analyses, reinforcement learning simulations, and hierarchical Bayesian model fitting. Across experiments, reward volatility exerted an inverse U-shaped nonlinear effect on the arbitration between model-free and model-based reinforcement learning strategies, as the model-based learning strategy was strongly driven at intermediate levels of reward volatility. These modulation effects were observed only in individuals who had learned the transition structure in the task, whereas those who had not learned the transition structure relied on the model-free learning strategy regardless of reward volatility. Reinforcement learning simulations revealed that the relative advantage of the model-based learning strategy over the model-free learning strategy peaked at intermediate levels of reward volatility. Additionally, increased time pressure shifted behavior toward the model-free learning strategy. These results demonstrated that, humans do not always use the model-based reinforcement learning strategy in uncertain and dynamic environments, even when they are aware of the task structure, supporting cost-benefit optimization. Author SummaryThe ability to flexibly guide behavior by carefully considering future consequences is fundamental to a prominent property of human intelligence and rationality. However, what drives this deliberative system? In this study, we investigated the factors that promote deliberative versus habitual behavior using decision-making tasks with uncertain structures and changing rewards. We found that participants who spontaneously learned the hidden transition structure in the task used this knowledge to guide deliberative behavior. Conversely, participants who did not learn the structure relied primarily on habitual strategies, repeating actions that had previously been rewarded. Among participants who learned the structure, the degree of deliberative behavior changed nonlinearly with reward volatility, in which the speed at which rewards changed over time. We also observed that limiting the decision time reduced deliberative behavior and promoted habitual responding. These findings suggest that under uncertain and dynamic environments, deliberative control is adaptively regulated according to cost-benefit optimization. Our results contribute to understanding how humans flexibly adjust their behavioral control systems in response to environmental conditions.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
PLOS Computational Biology
1863 papers in training set
Top 1%
21.7%
2
Scientific Reports
3612 papers in training set
Top 5%
9.6%
3
Psychological Review
19 papers in training set
Top 0.1%
7.8%
4
Royal Society Open Science
214 papers in training set
Top 0.2%
7.8%
5
Cognition
47 papers in training set
Top 0.1%
5.4%
50% of probability mass above
6
PLOS ONE
5266 papers in training set
Top 31%
4.8%
7
Proceedings of the Royal Society B: Biological Sciences
393 papers in training set
Top 2%
3.2%
8
eLife
5828 papers in training set
Top 35%
3.2%
9
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 18%
3.1%
10
Journal of Experimental Psychology: General
23 papers in training set
Top 0.1%
2.6%
11
iScience
1154 papers in training set
Top 11%
2.4%
12
Nature Communications
5641 papers in training set
Top 40%
2.4%
13
Journal of Neurophysiology
302 papers in training set
Top 2%
1.9%
14
Communications Psychology
22 papers in training set
Top 0.2%
1.7%
15
European Journal of Neuroscience
189 papers in training set
Top 2%
1.5%
16
Nature Human Behaviour
95 papers in training set
Top 1%
1.4%
17
Experimental Brain Research
53 papers in training set
Top 0.6%
1.1%
18
Frontiers in Neuroscience
256 papers in training set
Top 5%
1.1%
19
The Journal of Neuroscience
1025 papers in training set
Top 9%
1.1%
20
Journal of The Royal Society Interface
235 papers in training set
Top 3%
1.1%
21
Animal Cognition
23 papers in training set
Top 0.4%
1.1%
22
Philosophical Transactions of the Royal Society B
51 papers in training set
Top 0.9%
0.8%
23
Behavioral Neuroscience
25 papers in training set
Top 0.3%
0.8%
24
Behavioural Brain Research
77 papers in training set
Top 2%
0.8%
25
PeerJ
308 papers in training set
Top 11%
0.8%
26
Animal Behaviour
73 papers in training set
Top 1.0%
0.6%