Meta-learning is expressed through altered prefrontal cortical dynamics
Sun, X.; Comrie, A. E.; Kahn, A. E.; Monroe, E. J.; Joshi, A.; Guidera, J. A.; Denovellis, E. L.; Krausz, T. A.; Zhou, J.; Thompson, P.; Hernandez, J.; Yorita, A.; Haque, R.; Berke, J. D.; Daw, N. D.; Frank, L. M.
Show abstract
Learning where and when rewards like food and water are available is essential for survival1,2. In the simplest cases where resource availability is stable, animals can learn reward contingencies by integrating outcomes across repeated samples of each possible action. In more natural settings, however, reward availability is governed by structured higher-order rules such as depletion and repletion over time. To adapt flexibly to such changing environments, optimal choices require meta-learning wherein animals learn how to learn from external feedback, ultimately enabling them to infer the underlying reward structure from abstract, generalizable rules rather than relying solely on recent outcomes3,4. The existence of meta-learning in animal behavior is well established3-8, yet the neural circuits and computations that implement it remain poorly understood9-11. Here we investigated meta-learning using a spatial foraging task in which rats acquired a depletion-repletion rule that regulated reward availability, and carried out longitudinal, high-density recordings from the medial prefrontal cortex (mPFC). We show that meta-learning engages specific, systematic changes in mPFC neural dynamics that embed the learned rule and thereby alter how the network learns action values from reward outcomes. These dynamics are based on mixed coding of task structure and value in individual mPFC neurons. At the population level, this coding organizes into low-dimensional dynamical motifs that generalize across task conditions. As meta-learning progresses, these motifs are reshaped to instantiate both rule-guided inference of future states before outcome delivery and rule-based value updating during the outcome period. These results indicate that meta-learning sculpts pre-existing prefrontal dynamics to support the acquisition of new, generalizable reward-learning strategies.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.