Vector-valued dopamine improves learning of continuous outputs in the striatum
Waernberg, E.; Kumar, A.
Show abstract
It is well established that midbrain dopaminergic neurons support reinforcement learning (RL) in the basal ganglia by transmitting a reward prediction error (RPE) to the striatum. In particular, different computational models and experiments have shown that a striatumwide RPE signal can support RL over a small discrete set of actions (e.g. no/no-go, choose left/right). However, there is accumulating evidence that the basal ganglia functions not as a selector between predefined actions, but rather as a dynamical system with graded, continuous outputs. To reconcile this view with RL, there is a need to explain how dopamine could support learning of dynamic outputs, rather than discrete action values. Inspired by the recent observations that besides RPE, the firing rates of midbrain dopaminergic neurons correlate with motor and cognitive variables, we propose a model in which dopamine signal in the striatum carries a vector-valued error feedback signal (a loss gradient) instead of a homogeneous scalar error (a loss). Using a recurrent network model of the basal ganglia, we show that such a vector-valued feedback signal results in an increased capacity to learn a multidimensional series of real-valued outputs. The corticostriatal plasticity rule we employed is based on Random Feedback Learning Online learning and is a fully local, "three-factor" product of the presynaptic firing rate, a post-synaptic factor and the unique dopamine concentration perceived by each striatal neuron. Crucially, we demonstrate that under this plasticity rule, the improvement in learning does not require precise nigrostriatal synapses, but is compatible with random placement of varicosities and diffuse volume transmission of dopamine.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Direct-fit to nature: an evolutionary perspective on biological (and artificial) neural networks 95%
- Optimal anticipatory control as a theory of motor preparation: a thalamo-cortical circuit model 95%
- Neural trajectories in the supplementary motor area and primary motor cortex exhibit distinct geometries, compatible with different classes of computation 95%
Similar papers in this journal
- Orchestrated Excitatory and Inhibitory Learning Rules Lead to the Unsupervised Emergence of Self-sustained and Inhibition-stabilized Dynamics 97%
- Are place cells just memory cells? Memory compression leads to spatial tuning and history dependence 97%
- Bayesian inference in ring attractor networks 96%
Similar papers in this journal
- A solution to the learning dilemma for recurrent networks of spiking neurons 97%
- Latent Representations in Hippocampal Network Model Co-Evolve with Behavioral Exploration of Task Structure 96%
- Tonic dopamine and biases in value learning linked through a biologically inspired reinforcement learning model 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.