関連する実験動画
Updated: Jan 25, 2026

04:48
Control of Eating Behavior Using a Novel Feedback System
Published on: May 8, 2018
11.5K
割引逆強化学習を用いた線形連続時間システムの出力フィードバック制御
IEEE transactions on cybernetics
|January 23, 2026
まとめ
本研究では、出力データのみを使用して未知のシステムを制御するための新しい割引逆強化学習(DIRL)アルゴリズムを導入します。この手法は、状態を再構成し、最適な制御ポリシーを効率的に学習し、既存の手法よりも優れた性能を発揮します。
科学分野:
- 制御システム工学
- 機械学習
- ロボット工学
背景:
- 割引逆強化学習(DIRL)は通常、完全な状態フィードバックを必要とし、入力-出力データのみを持つ実世界のアプリケーションでの使用を制限します。
- 部分的に観測可能な状態を持つ未知の連続時間(CT)システムは、重大な制御上の課題をもたらします。
- 最適な制御ポリシー導出のために未知の割引値関数を学習することが重要です。
研究 の 目的:
- 未知のCTシステムの線形二次(LQ)制御のための新しいモデルフリー、出力フィードバック(OPFB)DIRLアルゴリズムを開発すること。
- 入力-出力データからの学習を可能にすることにより、既存のDIRL法の限界に対処すること。
- ポリシー学習のために専門家の制御出力データを利用してシステム状態を再構成すること。
主な方法:
- 専門家の制御と測定された出力データを利用して状態再構成法を設計しました。
- 未知の値関数と最適な制御ポリシーを反復的に学習するために、モデルフリーのOPFB DIRLアルゴリズムを提示しました。
- アルゴリズムの収束と解の一意性の厳密な分析を実行しました。
主要な成果:
- 提案されたアルゴリズムは、専門家の制御ポリシーを効果的に回復します。
- シミュレーションは、最先端の方法と比較して優れた計算効率を示しています。
- アルゴリズムは、部分的に観測可能な状態と未知の値関数を正常に処理します。
結論:
- 新しいOPFB DIRLアルゴリズムは、限られた状態情報を持つ未知のCTシステムを制御するための効果的なソリューションを提供します。
- この方法は、入力-出力データのみを利用することにより、実際的なシナリオでのDIRLの適用性を高めます。
- アルゴリズムは、最適な制御ポリシーを学習するための計算効率が高く堅牢なアプローチを提供します。
関連する概念動画
Feedback control systems
703
Feedback control systems are categorized in various ways based on their design, analysis, and signal types.
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
703
Linear time-invariant Systems
890
A system is linear if it displays the characteristics of homogeneity and additivity, together termed the superposition property. This principle is fundamental in all linear systems. Linear time-invariant (LTI) systems include systems with linear elements and constant parameters.
The input-output behavior of an LTI system can be fully defined by its response to an impulsive excitation at its input. Once this impulse response is known, the system's reaction to any other input can be...
The input-output behavior of an LTI system can be fully defined by its response to an impulsive excitation at its input. Once this impulse response is known, the system's reaction to any other input can be...
890
BIBO stability of continuous and discrete -time systems
898
System stability is a fundamental concept in signal processing, often assessed using convolution. For a system to be considered bounded-input bounded-output (BIBO) stable, any bounded input signal must produce a bounded output signal. A bounded input signal is one where the modulus does not exceed a certain constant at any point in time.
To determine the BIBO stability, the convolution integral is utilized when a bounded continuous-time input is applied to a Linear Time-Invariant (LTI) system....
To determine the BIBO stability, the convolution integral is utilized when a bounded continuous-time input is applied to a Linear Time-Invariant (LTI) system....
898
Linear Momentum in Control Volume
1.3K
Newton's second law is applied to obtain the linear momentum in a control volume in a fluid system. According to this law, the rate of change of linear momentum is equal to the sum of external forces acting on the system. When a control volume matches the fluid system at a specific moment, the forces acting on both are identical. Reynolds transport theorem helps explain this by breaking down the system's linear momentum into two components: the rate of change of linear momentum within...
1.3K
Root Loci for Positive-Feedback Systems
338
The Hartley oscillator is a positive feedback system that sustains oscillations by feeding the output back to the input in phase, thereby reinforcing the signal. Positive feedback systems can be viewed as negative feedback systems with inverted feedback signals. In these systems, the root locus encompasses all points on the s-plane where the angle of the system transfer function equals 360 degrees.
The construction rules for the root locus in positive feedback systems are similar to those in...
The construction rules for the root locus in positive feedback systems are similar to those in...
338
Control Systems
1.8K
Control systems are everywhere in contemporary society, influencing diverse applications from aerospace to automated manufacturing. These systems can be found naturally within biological processes, such as blood sugar regulation and heart rate adjustment in response to stress, as well as in man-made systems like elevators and automated vehicles. A control system is essentially a network of subsystems and processes that collaboratively convert specific inputs into desired outputs.
At the heart...
At the heart...
1.8K

