Optimal Reinforcement Learning with Asymmetric Updating in Volatile Environments: a Simulation Study
Rostami Kandroodi, M.; Vahabie, A.-h.; Ahmadi, S.; Nadjar Araabi, B.; Nili Ahmadabadi, M.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWThe ability to predict the future is essential for decision-making and interaction with the environment to avoid punishment and gain reward. Reinforcement learning algorithms provide a normative way for interactive learning, especially in volatile environments. The optimal strategy for the classic reinforcement learning model is to increase the learning rate as volatility increases. Inspired by optimistic bias in humans, an alternative reinforcement learning model has been developed by adding a punishment learning rate to the classic reinforcement learning model. In this study, we aim to 1) compare the performance of these two models in interaction with different environments, and 2) find optimal parameters for the models. Our simulations indicate that having two different learning rates for rewards and punishments increases performance in a volatile environment. Investigation of the optimal parameters shows that in almost all environments, having a higher reward learning rate compared to the punishment learning rate is beneficial for achieving higher performance which in this case is the accumulation of more rewards. Our results suggest that to achieve high performance, we need a shorter memory window for recent rewards and a longer memory window for punishments. This is consistent with optimistic bias in human behavior.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Collective Evolution Learning Model for Vision-Based Collective Motion with Collision Avoidance 95%
- Screening plans for SARS-CoV-2 based on sampling and rotation: an example in the school setting 94%
- A Hessian-based decomposition characterizes how performance in complex motor skills depends on individual strategy and variability 94%
Similar papers in this journal
- Episodic-Like Memory in a Simulation of Cuttlefish Behavior 95%
- Nonlinear neural network dynamics accounts for human confidence in a sequence of perceptual decisions 94%
- Ranking the Effectiveness of Non-Pharmaceutical Interventions to Counter COVID-19 in UK Universities with Vaccinated Population 93%
Similar papers in this journal
- A Complex-valued Oscillatory Neural Network for Storage and Retrieval of Multichannel Electroencephalogram Signals 93%
- A Multiscale, Systems-level, Neuropharmacological Model of Cortico-Basal Ganglia System for Arm Reaching under Normal, Parkinsonian and Levodopa Medication Conditions 92%
- A Network Architecture for Bidirectional Neurovascular Coupling in Rat Whisker Barrel Cortex 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.