Back

Dopamine reports reward prediction errors, but does not update policy, during inference-guided choice

Blanco-Pozo, M.; Akam, T.; Walton, M. E.

2021-06-27 neuroscience
10.1101/2021.06.25.449995 bioRxiv
Show abstract

Rewards are thought to influence future choices through dopaminergic reward prediction errors (RPEs) updating stored value estimates. However, accumulating evidence suggests that inference about hidden states of the environment may underlie much adaptive behaviour, and it is unclear how these two accounts of reward-guided decision-making should be integrated. Using a two-step task for mice, we show that dopamine reports RPEs using value information inferred from task structure knowledge, alongside information about recent reward rate and movement. Nonetheless, although rewards strongly influenced choices and dopamine, neither activating nor inhibiting dopamine neurons at trial outcome affected future choice. These data were recapitulated by a neural network model in which frontal cortex learned to track hidden task states by predicting observations, while basal ganglia learned corresponding values and actions via dopaminergic RPEs. Together, this two-process account reconciles how dopamine-independent state inference and dopamine-mediated reinforcement learning interact on different timescales to determine reward-guided choices.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.