Noradrenaline modulates tabula-rasa exploration
Dubois, M.; Habicht, J.; Michely, J.; Moran, R.; Dolan, R.; Hauser, T.
Show abstract
An exploration-exploitation trade-off, the arbitration between sampling a lesser-known against a known rich option, is thought to be solved using computationally demanding exploration algorithms. Given known limitations in human cognitive resources, we hypothesised the presence of additional cheaper strategies. We examined for such heuristics in choice behaviour where we show this involves a value-free random exploration, that ignores all prior knowledge, and a novelty exploration that targets novel options alone. In a double-blind, placebo-controlled drug study, assessing contributions of dopamine (400mg amisulpride) and noradrenaline (40mg propranolol), we show that value-free random exploration is attenuated under the influence of propranolol, but not under amisulpride. Our findings demonstrate that humans deploy distinct computationally cheap exploration strategies and where value-free random exploration is under noradrenergic control. Data and materials availabilityData and code will be provided upon acceptance.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Dopamine enhances model-free credit assignment through boosting of retrospective model-based inference 97%
- A circuit mechanism for irrationalities in decision-making and NMDA receptor hypofunction: behaviour, computational modelling, and pharmacology 96%
- Sex differences in learning from exploration 96%
Similar papers in this journal
- Combined model-free and model-sensitive reinforcement learning in non-human primates 97%
- Removal of reinforcement improves instrumental performance in humans by decreasing a general action bias rather than unmasking learnt associations 96%
- An inductive bias for slowly changing features in human reinforcement learning 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.