Related Experiment Video
Updated: Jun 5, 2025

A Lateralized Odor Learning Model in Neonatal Rats for Dissecting Neural Circuitry Underpinning Memory Formation
Published on: August 18, 2014
Revisiting the problem of learning long-term dependencies in recurrent neural networks
Liam Johnston1, Vivak Patel1, Yumian Cui1
1Department of Statistics, University of Wisconsin, Madison, WI, USA.
Recurrent neural networks (RNNs) can learn long-term dependencies despite the vanishing and exploding gradient (VEG) problem. Hyper-parameter tuning, especially learning rate, significantly impacts RNNs' ability to learn complex sequential data.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Deep Learning
Background:
- Recurrent Neural Networks (RNNs) are crucial for sequential data.
- The vanishing and exploding gradient (VEG) problem is widely believed to hinder RNNs' ability to learn long-term dependencies.
- This belief has driven significant research in RNN advancements.
Purpose of the Study:
- To challenge the established belief that VEG prevents RNNs from learning long-term dependencies.
- To investigate the actual learning dynamics of RNNs beyond the VEG and hyperbolic attractor explanations.
- To identify key factors influencing RNNs' capacity for long-term dependency learning.
Main Methods:
- Trained over 40,000 RNNs using a large factorial experiment.
- Re-examined the concept of latching behavior in RNNs.
- Analyzed the impact of hyper-parameters, specifically learning rate, on model performance.
Main Results:
- Provided evidence contradicting the long-held belief that VEG inherently prevents long-term dependency learning in RNNs.
- Demonstrated that RNNs can learn dynamics not fully explained by hyperbolic attractors.
- Identified learning rate as a critical hyper-parameter influencing the success of long-term dependency learning.
Conclusions:
- The VEG problem does not necessarily preclude RNNs from learning long-term dependencies.
- RNNs exhibit diverse learned dynamics beyond hyperbolic attractors.
- Optimizing hyper-parameters, particularly learning rate, is essential for enabling RNNs to effectively learn long-term dependencies.
Related Concept Videos
Long-term Potentiation
Higher Mental Functions of Brain: Learning and Memory
Long-term Depression
Calcium Ion Concentration Mechanism
If over...
Neural Circuits
Neuronal pools are collections of nerve cells with similar functions and interact through chemical and electrical signals. These pools include both interneurons (the central neural circuit nodes that...
Real-World Application of Classical Conditioning
Higher-order, or second-order, conditioning occurs when a neutral stimulus becomes associated with an already established conditioned stimulus through repeated pairings. For instance, if a dog has been...
Elaborative Rehearsals
The effectiveness of...

