Related Experiment Video
Updated: Jan 11, 2026

Measuring Attention and Visual Processing Speed by Model-based Analysis of Temporal-order Judgments
Published on: January 23, 2017
Ratio divergence learning using target energy in restricted Boltzmann machines: Beyond Kullback-Leibler divergence
Yuichi Ishida1, Yuma Ichikawa1,2, Aki Dote1
1Fujitsu Ltd., Kawasaki, Japan.
Ratio divergence (RD) learning enhances discrete energy-based models by combining forward and reverse Kullback-Leibler divergence (KLD) learning. This novel approach improves energy function fitting and learning stability, especially for complex models.
Area of Science:
- Machine Learning
- Artificial Intelligence
- Computational Statistics
Background:
- Discrete energy-based models (EBMs) are powerful tools for modeling complex data distributions.
- Restricted Boltzmann machines (RBMs) are a fundamental discrete EBM satisfying the universal approximation theorem.
- Existing learning methods like forward and reverse Kullback-Leibler divergence (KLD) learning have limitations, including underfitting and mode collapse.
Purpose of the Study:
- To introduce Ratio Divergence (RD) learning, a new method for discrete energy-based models.
- To address the limitations of traditional KLD learning methods in discrete EBMs.
- To evaluate the effectiveness of RD learning compared to existing methods.
Main Methods:
- RD learning utilizes both training data and a tractable target energy function.
- RD learning combines the strengths of forward and reverse KLD learning.
- Numerical experiments were conducted using various discrete EBMs, including RBMs, with RD learning as the primary method.
Main Results:
- RD learning significantly outperforms existing methods in energy function fitting, mode-covering, and learning stability.
- The effectiveness of RD learning increases with the dimensionality of the target models.
- A baseline method summing forward and reverse KLD was evaluated against RD learning.
Conclusions:
- RD learning offers a robust and effective approach for training discrete energy-based models.
- RD learning successfully mitigates underfitting and mode collapse issues inherent in KLD learning.
- The proposed RD learning method shows promise for complex, high-dimensional discrete distribution modeling.
Related Concept Videos
Maxwell-Boltzmann Distribution: Problem Solving
This distribution function f(v) is defined by saying that the expected number N (v1,v2) of particles with speeds between v1 and v2 is given by
Divergence and Stokes' Theorems
Energy Conservation and Bernoulli's Equation
All the terms in the equation have the dimension of energy per unit volume. The kinetic energy per unit volume is called the kinetic energy density, and the potential energy per unit volume is...
Conservation of Energy in Control Volume
For steady flow systems, the time derivative of the stored energy becomes zero since there is no energy accumulation within the control volume. This simplifies the energy equation to:
Difference from Background: Limit of Detection
The LOD indicates the presence or absence...
Conservation of Energy: Application