Back

Entropy regularised reinforcement learning reconciles aversive and action prediction errors in the tail of the striatum

Mahajan, P.; Seymour, B.

2026-07-11 neuroscience
10.64898/2026.07.09.737461 bioRxiv
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWDopamine activity in the tail of the striatum (TS) presents a novel challenge for reinforcement-learning theories of dopamine. Some studies suggest that TS-projecting dopamine signals encode aversive or threat prediction errors, whereas others argue that they encode action prediction errors involved in soft-habit formation. Here, we show that these accounts need not be mutually exclusive. We instantiate an entropy-regularised reinforcement-learning model in which TS-projecting dopamine neurons update both aversive values and the default policy. In this model, threat belief gates aversive value initialisations, producing TS-like activity during retreat from potentially threatening novel objects, while default-policy learning generates action prediction error signals that decline as actions become habitual. Our results further suggest why both of these signals may need to coexist in the temporal difference errors in our model, and qualitatively reproduce key response patterns and simulations from studies previously used to support both views. Beyond this descriptive reconciliation, our model simulations also highlight the normative role of the tail of the striatum in cautious behaviours in the context of potential threats and stable learning in the face of outcome uncertainty.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.