Behavioural and neural mechanisms for stochastic choices in mixed strategy games
Aloor, J.; Sit, T. P.; Gauld, O. M.; Warren, J.; Mower, M.; Lee, D.; Duan, C. A.
Show abstract
Adaptive behaviour usually requires exploiting regularities in the environment, but in competitive settings the opposite can be true: predictable choice patterns can be exploited by others, making unpredictability itself advantageous. How neural circuits generate such strategic variability remains poorly understood. Here, we trained mice to play a zero-sum game against an opponent that exploited statistical regularities in their choices and rewards, and tracked their behaviour and dorsal cortical dynamics across learning. Using a hidden Markov model, we found that mice transitioned from structured, predictable strategies towards a near-optimal stochastic strategy as they learned. Applying the same framework to monkeys playing the same game identified a shared stochastic strategy across species, despite differences in how it was deployed. Cortex-wide imaging revealed that stochastic choices were associated with reduced representation of reward history, while immediate reward signals remained robust. Critically, while the strength of cortical reward signals predicted subsequent choice during reward-guided behaviour, this relationship was abolished during stochastic behaviour. Thus, adaptive stochasticity does not simply arise from a loss of reward information, but from selectively decoupling reward from future choice. These results reveal a neural mechanism through which animals suppress otherwise useful reward-guided structure to generate adaptive unpredictability in competitive environments.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.