Reinforcement learning: computing the temporal difference of values via distinct corticostriatal pathways
Kenji Morita1, Mieko Morishima, Katsuyuki Sakai
1Physical and Health Education, Graduate School of Education, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo 113-0033, Japan. morita@p.u-tokyo.ac.jp
Trends in Neurosciences
|June 5, 2012
Summary
This study proposes a novel mechanism for how midbrain dopamine neurons compute reward prediction error signals. It involves distinct corticostriatal neuron subpopulations representing current and past states to calculate temporal differences, offering a unified view of basal ganglia function.
Area of Science:
- Neuroscience
- Computational Neuroscience
- Systems Neuroscience
Background:
- Midbrain dopamine neurons are implicated in reward prediction error (RPE) signaling.
- The precise neural mechanisms underlying RPE computation remain largely unknown.
Purpose of the Study:
- To propose a novel computational mechanism for RPE signaling in corticostriatal circuits.
- To explain how distinct neuronal subpopulations contribute to representing temporal differences in value.
Main Methods:
- Theoretical modeling based on recent findings in corticostriatal circuits.
- Hypothesizing specific connectivity patterns and neuronal dynamics.
Main Results:
- Two distinct corticostriatal neuron subpopulations differentially encode current and previous states/actions.
- Unidirectional connectivity and recurrent excitation within subpopulations are key.
- Selective connections to basal ganglia direct and indirect pathways enable temporal difference computation.
Conclusions:
- This model provides a unified framework for understanding basal ganglia functions.
- The proposed mechanism offers insights into the neural basis of RPE and has potential clinical implications for neurological and psychiatric disorders.
Related Concept Videos
Long-term Potentiation
Long-term potentiation, or LTP, is one of the ways by which synaptic plasticity—changes in the strength of chemical synapses—can occur in the brain. LTP is the process of synaptic strengthening that occurs over time between pre- and postsynaptic neuronal connections. The synaptic strengthening of LTP works in opposition to the synaptic weakening of long-term depression (LTD) and together are the main mechanisms that underlie learning and memory.
Long-term Potentiation
Long-term potentiation, or LTP, is one of the ways by which synaptic plasticity—changes in the strength of chemical synapses—can occur in the brain. LTP is the process of synaptic strengthening that occurs over time between pre and postsynaptic neuronal connections. The synaptic strengthening of LTP works in opposition to the synaptic weakening of long-term depression (LTD) and together are the main mechanisms that underlie learning and memory.
Hebbian LTP
LTP can occur when presynaptic neurons...
Hebbian LTP
LTP can occur when presynaptic neurons...


