Related Experiment Videos
A new criterion using information gain for action selection strategy in reinforcement learning
Kazunori Iwata1, Kazushi Ikeda, Hideaki Sakai
1Graduate School of Informatics, Kyoto University, Kyoto 606-8501, Japan. kiwata@sys.i.kyoto-u.ac.jp
IEEE Transactions on Neural Networks
|October 6, 2004
Summary
This study introduces a new w-based strategy for probabilistic action selection, outperforming the traditional Q-based approach by utilizing information gain for better predictions in financial return sequences.
Area of Science:
- Machine Learning
- Information Theory
- Quantitative Finance
Background:
- Financial return sequences can be modeled as outputs from a parametric compound source.
- The coding rate of a source quantifies the information it contains about returns.
Purpose of the Study:
- To develop novel l-learning algorithms for estimating expected information gain.
- To introduce a new criterion, the w-ratio, for probabilistic action-selection strategies.
- To evaluate the performance of the w-based strategy against conventional methods.
Main Methods:
- Parametric compound source modeling of return sequences.
- Predictive coding for estimating information gain.
- Development of l-learning algorithms.
- Introduction of the w-ratio (return loss to information gain).
Main Results:
- Convergence proof for the estimated information gain.
- The w-based strategy demonstrated superior performance compared to the Q-based strategy in experiments.
- The w-ratio effectively balances return loss and information gain.
Conclusions:
- The proposed l-learning algorithms and w-ratio offer a promising new direction for probabilistic action-selection.
- Information gain is a valuable metric for improving decision-making in financial contexts.
- The w-based strategy provides a robust alternative to existing methods like Q-learning.