Back

Optimal Practice Schedules in a Dual-Rate Model of Motor Adaptation, and Their Recovery by Reinforcement Learning

Jeter, R.; Todorov, D.; Molkov, Y.

2026-06-22 neuroscience
10.64898/2026.06.17.732970 bioRxiv
Show abstract

A clinician guiding a stroke patient through a 45-minute rehabilitation session, a coach planning a training day, a teacher choosing the order of practice problems, they all face the same question: "given everything practiced so far, what should the next trial be?" The motor-learning literature offers two coarse answers, blocked and interleaved ("random") practice, with a well-known dissociation, blocked practice gives faster acquisition but worse retention, while interleaved practice gives the opposite. We argue that this dissociation is not a fixed property of practice schedules but a shadow of a richer structure. In particular, for a learner whose memory has a fast shared component and slower context-specific components, the best schedule should be a function of the learners current internal state and the time remaining before the retention probe. We make this precise in a minimal two-context fast-slow learner model whose optimal schedules can be computed exactly for short sessions and approximated by a structured beam-search upper bound for longer ones. The optimal schedule is not blocked, not interleaved, and not a single rule; it is a family of schedules determined by how much retention is weighted relative to acquisition. The family has three regimes (alternating, mixed, blocked-with-late-correction) and for long sessions, the optimal schedule has an interpretable structure -- exploit one context, repair the neglected one, then interleave to lock in retention. We then investigate whether a reinforcement-learning teacher, observing only the learners actions and errors without access to their internal memory states, can learn these optimal policies from interaction alone. Comparing these learned policies against the exact optima, we show that a model-free agent (PPO) recovers the short-horizon schedules and the long-horizon block-repair-interleave motif in the intermediate regime, but the benchmark also exposes a sharp failure in the acquisition-dominated regime, where PPO collapses to pure blocking and misses a sparse terminal correction. A warm-start diagnostic shows this failure is a genuine metastability of policy gradients rather than a tuning artifact, with blocked-plus-switch and pure-blocked acting as competing attractors that PPO cannot stabilize between. A hyperparameter sweep over observation history reveals that the agent requires very little behavioral context to plan optimally, demonstrating that partial observability is not a major barrier to finding optimal practice schedules. Finally, we discuss the implications of our framework for motor adaptation and contextual interference, offering practical insights on how instructors can design finite practice sessions to favor long-term retention.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

1
PLOS Computational Biology
1863 papers in training set
Top 1%
18.6%
2
Journal of Neurophysiology
302 papers in training set
Top 0.5%
7.9%
3
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 6%
7.3%
4
eLife
5828 papers in training set
Top 20%
5.6%
5
Nature Communications
5641 papers in training set
Top 27%
5.5%
6
Nature Neuroscience
252 papers in training set
Top 2%
4.4%
7
Journal of The Royal Society Interface
235 papers in training set
Top 0.9%
4.4%
50% of probability mass above
8
Scientific Reports
3612 papers in training set
Top 30%
3.4%
9
Neural Computation
39 papers in training set
Top 0.3%
3.3%
10
Communications Psychology
22 papers in training set
Top 0.1%
2.8%
11
Nature Human Behaviour
95 papers in training set
Top 0.7%
2.7%
12
Neuron
337 papers in training set
Top 3%
2.4%
13
PLOS ONE
5266 papers in training set
Top 45%
2.1%
14
Biological Cybernetics
15 papers in training set
Top 0.1%
1.5%
15
Frontiers in Computational Neuroscience
60 papers in training set
Top 0.8%
1.5%
16
The Journal of Neuroscience
1025 papers in training set
Top 8%
1.4%
17
Mathematical Biosciences
49 papers in training set
Top 0.8%
1.4%
18
Psychological Review
19 papers in training set
Top 0.2%
1.0%
19
iScience
1154 papers in training set
Top 31%
1.0%
20
Computational Psychiatry
12 papers in training set
Top 0.1%
1.0%
21
Journal of Vision
110 papers in training set
Top 0.7%
0.9%
22
Nature
645 papers in training set
Top 10%
0.8%
23
Frontiers in Artificial Intelligence
20 papers in training set
Top 0.7%
0.8%
24
eneuro
439 papers in training set
Top 8%
0.8%
25
Cell Systems
201 papers in training set
Top 4%
0.8%
26
NeuroImage
903 papers in training set
Top 6%
0.6%
27
PNAS Nexus
159 papers in training set
Top 4%
0.6%
28
Science Advances
1243 papers in training set
Top 33%
0.6%