部分観測を伴う有限ホライゾンH∞追跡のためのQ学習アプローチ
IEEE transactions on cybernetics
|January 19, 2026
まとめ
本研究は、部分観測を伴う離散時間システムのための新しいモデルフリー強化学習アルゴリズムを導入する。これらのデータ駆動型手法は、初期ポリシーや割引係数を必要とせずに、有限ホライゾンH∞追跡制御の課題に対処する。
科学分野:
- 制御理論
- 強化学習
- ゲーム理論
背景:
- 既存の強化学習(RL)手法は、多くの場合、完全な状態情報を必要とし、無限ホライゾン、時間不変システムに限定されます。
- 部分観測と未知のダイナミクスを持つ有限ホライゾン制御は、時間変化リッカチ方程式の必要性を含む重大な課題を提示します。
- モデルフリーアプローチは、ダイナミクスが未知のシステムで望ましく、入出力データのみに依存します。
研究 の 目的:
- 部分観測と未知のダイナミクスを持つ離散時間線形システムに対する有限ホライゾンH∞追跡制御問題を調査すること。
- 既存のアプローチの限界、特に状態情報とシステムホライゾンに関する限界を克服するモデルフリー強化学習アルゴリズムを開発すること。
- 初期許容ポリシーまたは割引係数を必要とせずに、時間変化制御問題のフレームワークを提供すること。
主な方法:
- データ駆動型システム表現を作成するために、過去の入出力軌道からシステム状態を再構築すること。
- 入出力データに基づいた時間変化Q関数の定義。
- モデルフリー、データ駆動型制御のために設計された2つのミニマックスQ学習アルゴリズムの提案。
主要な成果:
- 開発されたアルゴリズムは、システム状態を正常に再構築し、入出力ベースのQ関数を定義します。
- ミニマックスQ学習アルゴリズムは、初期許容ポリシーを必要とせず、割引係数を回避するため、安定性保証が向上します。
- このフレームワークは、構造変更なしで無限ホライゾンおよび時間変化システムに拡張可能であることを示しています。
結論:
- 提案されたデータ駆動型、モデルフリー強化学習アルゴリズムは、部分観測を伴う離散時間システムの有限ホライゾンH∞追跡制御問題を効果的に解決します。
- 理論的収束が証明され、シミュレーション結果がアルゴリズムの有効性を検証します。
- この研究は、特に未知のダイナミクスと部分状態情報を持つシステムにおいて、制御のための強化学習における重要な進歩を提供します。
関連する概念動画
Observational Learning
853
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
853
Schwarzschild Radius and Event Horizon
2.6K
No object with a finite mass can travel faster than the speed of light in a vacuum. This fact has an interesting consequence in the domain of extremely high gravitational fields.
The minimum speed required to launch a projectile from the surface of an object to which it is gravitationally bound so that it eventually escapes the object’s gravitational field is called the escape velocity. The escape velocity is independent of the mass of the object. Merging the idea of escape...
The minimum speed required to launch a projectile from the surface of an object to which it is gravitationally bound so that it eventually escapes the object’s gravitational field is called the escape velocity. The escape velocity is independent of the mass of the object. Merging the idea of escape...
2.6K
Naturalistic Observations
17.0K
If you want to understand how behavior occurs, one of the best ways to gain information is to simply observe the behavior in its natural context. However, people might change their behavior in unexpected ways if they know they are being observed. How do researchers obtain accurate information when people tend to hide their natural behavior? As an example, imagine that your professor asks everyone in your class to raise their hand if they always wash their hands after using the restroom. Chances...
17.0K
Conservation of Mass in Finite Cotrol Volume
1.8K
The principle of conservation of mass is a fundamental law in fluid mechanics and is applied using the continuity equation. We apply the concept to a finite control volume to derive the continuity equation.
A system is defined as a collection of unchanging contents, and the conservation of mass states that a system's mass is constant.
A system is defined as a collection of unchanging contents, and the conservation of mass states that a system's mass is constant.
1.8K
Mixtures of Gases: Dalton's Law of Partial Pressures and Mole Fractions
43.7K
Unless individual gases chemically react with each other, the individual gases in a mixture of gases do not affect each other’s pressure. Each gas in a mixture exerts the same pressure that it would exert if it were present alone in the container. The pressure exerted by each individual gas in a mixture is called its partial pressure.
43.7K
Avoidance Learning and Learned Helplessness
2.5K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
2.5K


