Neural basis of compositional control.

Naturalistic goal-directed behaviour often involves continuous actions directed at dynamically changing goals 1-3 . Just as microeconomics serves as a rigorous foundation for discrete choices, control theory can serve as a foundation for understanding choice in continuous ones 3,4 . In continuous contexts, behaviour is composed of blends of goals, and the closest analogue to choice is a strategic reweighting of goal-specific control policies 5,6 . Here, to understand the algorithmic and neural b
Naturalistic goal-directed behaviour often involves continuous actions directed at dynamically changing goals 1-3 . Just as microeconomics serves as a rigorous foundation for discrete choices, control theory can serve as a foundation for understanding choice in continuous ones 3,4 . In continuous contexts, behaviour is composed of blends of goals, and the closest analogue to choice is a strategic reweighting of goal-specific control policies 5,6 . Here, to understand the algorithmic and neural bases of continuous choice, we examined behaviour and brain activity in humans performing a continuous prey-pursuit task 7 . Using a newly developed control-theoretic decomposition of behaviour, we find that pursuit strategies are well described by a meta-controller dictating a mixture of lower-level controllers, each linked to specific pursuit goals. Neurons in the anterior cingulate cortex predict major changes in policy blends, whereas hippocampal neurons encode and update the latent policy state supporting early planning. Meanwhile, orbitofrontal cortex activity is consistent with an encoding of the current value structure of the task, rather than policy switching. Together these results are consistent with a tripartite functional division in which hippocampus serves as a state-estimating controller, anterior cingulate cortex serves as a meta-controller, and orbitofrontal cortex provides a value context signal.




