An anatomical substrate of credit assignment in reinforcement learning
Kornfeld, J. M.; Januszewski, M.; Schubert, P. J.; Jain, V.; Denk, W.; Fee, M. S.
Show abstract
A key problem in learning is credit assignment. Biological systems lack a plausible mechanism to implement the backpropagation approach, a method that underlies much of the dramatic progress in artificial intelligence. Here, we use automated connectomic analysis to show that the synaptic architecture of songbird basal ganglia (Area X) supports local credit assignment using a variant of a node perturbation algorithm proposed in a model of reinforcement learning. Using two volume electron microscopy (vEM) datasets, we find that key predictions of the model hold true: axons that encode exploratory variability terminate predominantly on dendritic shafts, while axons that encode song timing (context) terminate predominantly on spines. Based on the detailed EM data, we then built a biophysical model of reinforcement learning that suggests that the synaptic dichotomy between variability and context encoding axons facilitates efficient learning. In combination, these findings provide strong evidence for a general, biologically plausible credit assignment model in vertebrate basal ganglia learning. One Sentence SummaryUsing automated connectomic analysis and biophysical modeling, we show how the basal ganglia could solve the credit assignment problem on the synaptic level.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Ultrafast Simulation of Large-Scale Neocortical Microcircuitry with Biophysically Realistic Neurons 97%
- Learning prediction error neurons in a canonical interneuron circuit 97%
- The interplay between homeostatic synaptic scaling and homeostatic structural plasticity maintains the robust firing rate of neural networks 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.