A computational principle of habit formation
Lakshminarasimhan, K. J.
Show abstract
1Actions are influenced by multiple decision-making systems - including a goal-directed system that favors rewarded actions and a habit system that repeats past actions - but precisely when one system prevails is not known. We show that when competition between these systems is resolved by a winner-take-all mechanism, the precise condition for the emergence of habits can be cast in terms of the well-known probability matching principle. The theory embodies a trade-off in which exploitation, or overmatching, maximizes reward but strengthens habits, while paradoxically, exploration preserves goal-directed behavior by sacrificing rewards. This tradeoff can be averted if learning operates on abstract latent state representations whereby knowing the broader context allows for switching between two habits instead of avoiding one, thus maximizing rewards without forfeiting flexibility. The theory explains a range of animal behaviors as well as task-dependent effects of striatal manipulation, and suggests that neural mechanisms governing exploration implicitly control arbitration between decision-making systems.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Neural mechanisms of distributed value representations and learning strategies 96%
- Tonic dopamine and biases in value learning linked through a biologically inspired reinforcement learning model 96%
- Latent Representations in Hippocampal Network Model Co-Evolve with Behavioral Exploration of Task Structure 96%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.