Back

Experience resetting in reinforcement learning facilitates exploration-exploitation transitions during a behavioral task for primates

Sakamoto, K.; Okuzaki, H.; Sato, A.; Mushiake, H.

2021-10-02 animal behavior and cognition
10.1101/2021.09.30.462676 bioRxiv
Show abstract

The exploration-exploitation trade-off is a fundamental problem in re-inforcement learning. To study the neural mechanisms involved in this problem, a target search task in which exploration and exploitation phases appear alternately is useful. Monkeys well trained in this task clearly understand that they have entered the exploratory phase and quickly acquire new experiences by resetting their previous experiences. In this study, we used a simple model to show that experience resetting in the exploratory phase improves performance rather than decreasing the greediness of action selection, and we then present a neural network-type model enabling experience resetting.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.