Back

Dopamine drives a positive reward bias on human reinforcement learning

zalta, a.; Skvortsova, V.; Hewitt, S. R.; Moutoussis, M.; Nour, M. M.; Dolan, R. J.; Findling, C.; Hauser, T. U.; Wyart, V.

2025-12-26 neuroscience
10.64898/2025.12.24.696333 bioRxiv
Show abstract

Formal theories of reinforcement learning (RL) prescribe a clearly defined function for dopamine, namely modulating learning via reward prediction errors (RPEs). Yet, empirical evidence in humans remains scarce, and recent advances introducing noisy RL cast doubt on a simple one-to-one mapping between neurotransmitters and computational mechanisms. Here, we detail a double-blind, placebo-controlled, randomised pharmacological study using the dopamine precursor L-DOPA, while healthy volunteers performed a volatile two-armed bandit task. Behaviourally, L-DOPA decreased switching behaviour following below-average rewards. Algorithmic RL modelling of human behaviour supported a dual effect of L-DOPA on the rate and precision of learning. By leveraging recurrent neural networks (RNNs) as implementational models of RL, we explain this dual effect through a single inference-time modulation, whereby L-DOPA triggers a positive reward bias at the input of the recurrent layer that implements RL. Our findings highlight a unifying mechanism at the implementation level that explain seemingly disparate algorithmic effects of dopamine.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.