A Cortical Microcircuit for Region-Specific Credit Assignment in Reinforcement Learning
Chevy, Q.; Szadai, Z.; Hertäg, L.; Moll, M.; Gibson, E. T.; Costa, R. P.; Kepecs, A.
Show abstract
The distributed architecture of the cortex poses a fundamental challenge for reinforcement learning: how to assign credit specifically to regions that contribute to successful behavior? Cortical neurons can be driven by both global reinforcers, like rewards, and local sensory features, making it difficult to disentangle these influences. To address this, we investigated cortical reinforcement learning by manipulating the reward-predictive sensory modality during learning tasks, while monitoring key regulators of cortical activity--local inhibitory neurons, and cholinergic inputs. We found that VIP interneurons are broadly recruited by reward-predictive cues via a modality-independent cholinergic signal. However, when task demands aligned with local computation, SST interneurons suppressed VIP recruitment through an inhibitory feedback loop. A computational model demonstrates that this cholinergic-VIP-SST interneuron circuit motif enables targeted reinforcement learning and region-specific credit assignment in the cortex. These results offer a neurobiologically-grounded framework for how the cortex uses global reinforcement signals to direct plasticity to task-relevant regions, enabling those regions to adapt and fine-tune their responses.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Computational functions of precisely balanced neuronal assemblies in an olfactory memory network 98%
- Stimulus information guides the emergence of behavior related signals in primary somatosensory cortex during learning 98%
- Behavioral state and stimulus strength regulate the role of somatostatin interneurons in stabilizing network activity 98%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.