Back

A meta reinforcement learning account of behavioral adaptation to volatility in recurrent neural networks

Tuzsus, D.; Brands, A. M.; Pappas, I.; Peters, J.

2024-05-09 neuroscience
10.1101/2024.05.09.593363 bioRxiv
Show abstract

Natural environments often exhibit various degrees of volatility, ranging from slowly changing to rapidly changing contingencies. How learners adapt to changing environments is a central issue in both reinforcement learning theory and psychology. For example, learners may adapt to changes in volatility by increasing learning if volatility increases, and reducing it if volatility decreases. Computational neuroscience and neural network modeling work suggests that this adaptive capacity may result from a meta-reinforcement learning process (implemented for example in the prefrontal cortex), where past experience endows the system with the capacity to rapidly adapt to environmental conditions in new situations. Here we provide direct evidence for a meta-reinforcement learning account of adaptation to environmental volatility. Recurrent neural networks (RNNs) were trained on a restless four-armed bandit reinforcement learning problem under three different training regimes (low volatility training only, medium volatility training only, or meta-volatility training across a range of volatilities). We show that, in contrast to RNNs trained in the low volatility regime, RNNs trained under the medium or meta-volatility regimes exhibited a superior adaption to environmental volatility. This was reflected in a better performance, and computational modeling of the networks behavior revealed a more adaptive adjustment of learning and exploration to varying levels of volatility during test. Results extend the meta-RL account to volatility adaptation and confirm past experience as a crucial factor of adaptive decision-making.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.