A novel critic signal in identified midbrain dopaminergic neurons of mice training in operant tasks
Matsumoto, J.; Oberto, V.; Pompili, M. N.; Todorova, R.; Papaleo, F.; Nishijo, H.; Venance, L.; Vandecasteele, M.; Wiener, S. I.
Show abstract
In the canonical interpretation of dopaminergic neuron activity during Pavlovian conditioning, initially cell firing is triggered by unexpected rewards. Upon learning, activation instead follows the reward-predictive conditioned stimulus, and when expected rewards are withheld, firing is inhibited. However, little is known about dopaminergic neuron activity during the actual learning process in complex operant tasks. Here, we recorded optogenetically identified dopaminergic neurons of ventral tegmental area (VTA) in mice training in multiple, successive operant sensory discrimination tasks. A delay between nose-poke choices and trial outcome signals (for reward or punishment) probed for predictive activity. During training, but prior to criterion performance, firing rates signaled correct versus incorrect choices, but prior to outcome signals. Thus, the neurons predicted whether choices would be rewarded, despite the animals subthreshold behavioral performance. Surprisingly, these neurons also fired after reward delivery, as if the rewards had been unexpected according to the canonical view, but activity was inhibited after punishment signals, as if the reward had been expected after all. These inconsistencies suggest revision of theoretical formulations of dopaminergic neuronal activity to embody multiple roles in temporal difference learning and actor-critic models. Furthermore, on training trials when these neurons predicted that a given choice was correct and would be rewarded, surprisingly, the mice adhered to other non-rewarded and untrained task strategies (e.g., spatial alternation). The DA neurons reward prediction activity could serve as critic signals for the choices just made. This consistent with the notion that the brain must reconcile multiple Bayesian belief representations during learning. Significance statementThe canonical view of dopaminergic function based on classical conditioning studies evokes reward-prediction error (RPE) signaling. Here, in mice performing a series of novel operant tasks with a delay between behavioral responses and reward/punishment signals, some neurons fired differentially after correct vs incorrect responses, but prior to the trial outcome (reward/punishment) signal. Nevertheless, the animals performed at chance levels, employing behavioral strategies other than the one signaled by these neurons. Furthermore, these same neurons showed canonical RPE responses, increased firing after reward signals (typically interpreted as the reward being unexpected) and firing rate decreased with punishment signals (interpreted as the reward having been expected). These findings indicate that dopaminergic neurons can participate in diverse functions underlying learning different behavioral strategies.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Ventral pallidal GABAergic neuron calcium activity encodes cue-driven reward-seeking and persists in the absence of reward delivery 96%
- Dissociable Roles of Pallidal Neuron Subtypes in Regulating Motor Patterns 96%
- Prefrontal neural ensembles develop selective code for stimulus associations within minutes of novel experiences. 96%
Similar papers in this journal
- Distracting stimuli evoke ventral tegmental area responses in rats during ongoing saccharin consumption 96%
- Functional maturation and experience-dependent plasticity in adult-born olfactory bulb dopaminergic neurons 95%
- Brief sensory deprivation triggers plasticity of dopamine-synthesising enzyme expression in genetically labelled olfactory bulb dopaminergic neurons 95%
Similar papers in this journal
- Frontal-sensory cortical projections become dispensable for attentional performance upon a reduction of task demand in mice 95%
- The differential effect of optogenetic serotonergic manipulation on sustained motor actions and stationary waiting for future rewards in mice 95%
- Lateral Hypothalamic GABAergic neurons encode and potentiate sucrose palatability 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.