Experience resetting in reinforcement learning facilitates exploration-exploitation transitions during a behavioral task for primates
Sakamoto, K.; Okuzaki, H.; Sato, A.; Mushiake, H.
Show abstract
The exploration-exploitation trade-off is a fundamental problem in re-inforcement learning. To study the neural mechanisms involved in this problem, a target search task in which exploration and exploitation phases appear alternately is useful. Monkeys well trained in this task clearly understand that they have entered the exploratory phase and quickly acquire new experiences by resetting their previous experiences. In this study, we used a simple model to show that experience resetting in the exploratory phase improves performance rather than decreasing the greediness of action selection, and we then present a neural network-type model enabling experience resetting.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Opponent Learning with Different Representations in the Cortico-Basal Ganglia Pathways Can Develop Obsession-Compulsion Cycle 97%
- Sleep prevents catastrophic forgetting in spiking neural networks by forming joint synaptic weight representations 96%
- A spiking neural-network model of goal-directed behaviour 96%
Similar papers in this journal
- Predicting human decision making in psychological tasks with recurrent neural networks 96%
- Collective Evolution Learning Model for Vision-Based Collective Motion with Collision Avoidance 96%
- Training a spiking neuronal network model of visual-motor cortex to play a virtual racket-ball game using reinforcement learning 95%
Similar papers in this journal
- Inhibitory stabilized network behaviour in a balanced neural mass model of a cortical column 93%
- Synaptic pruning facilitates online Bayesian model selection 93%
- A novel density-based neural mass model for simulating neuronal network dynamics with conductance-based synapses and membrane current adaptation 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.