Dopamine drives a positive reward bias on human reinforcement learning
zalta, a.; Skvortsova, V.; Hewitt, S. R.; Moutoussis, M.; Nour, M. M.; Dolan, R. J.; Findling, C.; Hauser, T. U.; Wyart, V.
Show abstract
Formal theories of reinforcement learning (RL) prescribe a clearly defined function for dopamine, namely modulating learning via reward prediction errors (RPEs). Yet, empirical evidence in humans remains scarce, and recent advances introducing noisy RL cast doubt on a simple one-to-one mapping between neurotransmitters and computational mechanisms. Here, we detail a double-blind, placebo-controlled, randomised pharmacological study using the dopamine precursor L-DOPA, while healthy volunteers performed a volatile two-armed bandit task. Behaviourally, L-DOPA decreased switching behaviour following below-average rewards. Algorithmic RL modelling of human behaviour supported a dual effect of L-DOPA on the rate and precision of learning. By leveraging recurrent neural networks (RNNs) as implementational models of RL, we explain this dual effect through a single inference-time modulation, whereby L-DOPA triggers a positive reward bias at the input of the recurrent layer that implements RL. Our findings highlight a unifying mechanism at the implementation level that explain seemingly disparate algorithmic effects of dopamine.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- A circuit mechanism for irrationalities in decision-making and NMDA receptor hypofunction: behaviour, computational modelling, and pharmacology 97%
- Dopamine enhances model-free credit assignment through boosting of retrospective model-based inference 97%
- Clarifying the role of an unavailable distractor in human multiattribute choice 95%
Similar papers in this journal
- Dopamine and temporal discounting: revisiting pharmacology and individual differences 97%
- Neural evidence for boundary updating as the source of the repulsive bias in classification 95%
- Striatal Gradient in Value-Decay Explains Regional Differences in Dopamine Patterns and Reinforcement Learning Computations 94%
Similar papers in this journal
- Meta-Reinforcement Learning reconciles surprise, value and control in the anterior cingulate cortex. 96%
- The drift diffusion model as the choice rule in inter-temporal and risky choice: a case study in medial orbitofrontal cortex lesion patients and controls. 96%
- An inductive bias for slowly changing features in human reinforcement learning 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.