A hardwired neural circuit for temporal difference learning
Campbell, M. G.; Ra, Y.; Chen, Z.; Xu, S.; Burrell, M.; Matias, S.; Watabe-Uchida, M.; Uchida, N.
Show abstract
The neurotransmitter dopamine plays a major role in learning by acting as a teaching signal to update the brains predictions about rewards. A leading theory proposes that this process is analogous to a reinforcement learning algorithm called temporal difference (TD) learning, and that dopamine acts as the error term within the TD algorithm (TD error). Although many studies have demonstrated similarities between dopamine activity and TD errors1-5, the mechanistic basis for dopaminergic TD learning remains unknown. Here, we combined large-scale neural recordings with patterned optogenetic stimulation to examine whether and how the key steps in TD learning are accomplished by the circuitry connecting dopamine neurons and their targets. Replacing natural rewards with optogenetic stimulation of dopamine axons in the nucleus accumbens (NAc) in a classical conditioning task gradually generated TD error-like activity patterns in dopamine neurons by specifically modifying the task-related activity of NAc neurons expressing the D1 dopamine receptor (D1 neurons). In turn, patterned optogenetic stimulation of NAc D1 neurons in naive animals drove dopamine neuron spiking according to the TD error of the stimulation pattern, indicating that TD computations are hardwired into this circuit. The transformation from D1 neurons to dopamine neurons could be described by a biphasic linear filter, with a rapid positive and delayed negative phase, that effectively computes a temporal difference. This finding suggests that the time horizon over which the TD algorithm operates--the temporal discount factor--is set by the balance of the positive and negative components of the linear filter, pointing to a circuit-level mechanism for temporal discounting. These results provide a new conceptual framework for understanding how the computations and parameters governing animal learning arise from neurobiological components.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Unexpected contributions of striatal projection neurons coexpressing dopamine D1 and D2 receptors in balancing motor control 99%
- Individual differences in decision-making shape how mesolimbic dopamine regulates choice confidence and change-of-mind 99%
- Opponent control of behavior by dorsomedial striatal pathways depends on task demands and internal state 99%
Similar papers in this journal
- Acquisition of non-olfactory encoding improves odour discrimination in olfactory cortex 99%
- Fast updating feedback from piriform cortex to the olfactory bulb relays multimodal reward contingency signals during rule-reversal 98%
- Prediction error signals in anterior cingulate cortex drive task-switching 98%
Similar papers in this journal
- A cortical circuit mechanism for coding and updating task structural knowledge in inference-based decision-making 99%
- Ventral frontostriatal circuitry mediates the computation of reinforcement from symbolic gains and losses 98%
- Potentiation of active locomotor state by spinal-projecting serotonergic neurons 98%
Similar papers in this journal
- Slowly evolving dopaminergic activity modulates the moment-to-moment probability of movement initiation. 99%
- Layer 6 ensembles can selectively regulate the behavioral impact and layer-specific representation of sensory deviants 98%
- Mating activates neuroendocrine pathways signaling hunger in Drosophila females 98%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.