Back

Comment on 'Optimal Policy for Multi-Alternative Decisions'

Marshall, J. A. R.

2019-12-24 animal behavior and cognition
10.1101/2019.12.18.880872 bioRxiv
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWOptimality analysis has recently been proposed for value-based decision-making, in which decision agents are rewarded by the value of the selected option. This contrasts with psychophysics where decision agents are typically rewarded only if they choose the correct or best option. The analysis of optimal policies for value-based decisions raises interesting and surprising parallels with decision rules proposed for accuracy-based decisions in binary and multi-alternative cases, and explains experimentally-observed deviations from rationality. However, the analysis assumes that decision agents should treat time as a linear cost, and thus optimise their Bayes Risk from decisions. A more naturalistic assumption is that future rewards are geometrically discounted, since they are less likely to be realised in an uncertain world. Changing the way in which time is costed leads to substantive changes in the resulting optimal policies, explains empirical data that previously could not be explained, and makes falsifiable predictions for future experiments.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.