Imitation as a model-free process in human reinforcement learning
Najar, A.; Bonnet, E.; Bahrami, B.; Palminteri, S.
Show abstract
While there is not doubt that social signals affect human reinforcement learning, there is still no consensus about their exact computational implementation. To address this issue, we compared three hypotheses about the algorithmic implementation of imitation in human reinforcement learning. A first hypothesis, decision biasing, postulates that imitation consists in transiently biasing the learners action selection without affecting her value function. According to the second hypothesis, model-based imitation, the learner infers the demonstrators value function through inverse reinforcement learning and uses it for action selection. Finally, according to the third hypothesis, value shaping, demonstrators actions directly affect the learners value function. We tested these three psychologically plausible hypotheses in two separate experiments (N = 24 and N = 44) featuring a new variant of a social reinforcement learning task, where we manipulated the quantity and the quality of the demonstrators choices. We show through model comparison that value shaping is favored, which provides a new perspective on how imitation is integrated into human reinforcement learning.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Combined model-free and model-sensitive reinforcement learning in non-human primates 97%
- Removal of reinforcement improves instrumental performance in humans by decreasing a general action bias rather than unmasking learnt associations 97%
- Dynamic integration of forward planning and heuristic preferences during multiple goal pursuit 97%
Similar papers in this journal
- Novel entropy-based metrics for predicting choice behavior based on local response to reward 96%
- Controllability Governs the Balance Between Pavlovian and Instrumental Action Selection 95%
- Linear reinforcement learning: Flexible reuse of computation in planning, grid fields, and cognitive control 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.