Video Experimental Relacionado
Updated: Jan 25, 2026

04:48
Control of Eating Behavior Using a Novel Feedback System
Published on: May 8, 2018
11.5K
Control de Retroalimentación de Salida de Sistemas Lineales de Tiempo Continuo Utilizando Aprendizaje por Refuerzo
IEEE transactions on cybernetics
|January 23, 2026
Resumen
Este estudio presenta un nuevo algoritmo de aprendizaje por refuerzo inverso descontado (DIRL) para controlar sistemas desconocidos utilizando solo datos de salida. El método reconstruye estados y aprende políticas de control óptimas de manera eficiente, superando a las técnicas existentes.
Área de la Ciencia:
- Ingeniería de Sistemas de Control
- Aprendizaje Automático
- Robótica
Sus antecedentes:
- El aprendizaje por refuerzo inverso descontado (DIRL) típicamente requiere retroalimentación de estado completo, lo que limita su uso en aplicaciones del mundo real con solo datos de entrada-salida.
- Los sistemas de tiempo continuo (CT) desconocidos con estados parcialmente observables presentan desafíos de control significativos.
- El aprendizaje de funciones de valor descontadas desconocidas es crucial para la derivación de políticas de control óptimas.
Objetivo del estudio:
- Desarrollar un novedoso algoritmo DIRL de retroalimentación de salida (OPFB) sin modelo para el control lineal cuadrático (LQ) de sistemas CT desconocidos.
- Abordar las limitaciones de los métodos DIRL existentes permitiendo el aprendizaje a partir de datos de entrada-salida.
- Reconstruir los estados del sistema utilizando datos de salida de control experto para el aprendizaje de políticas.
Principales métodos:
- Se diseña un método de reconstrucción de estado utilizando datos de control experto y datos de salida medidos.
- Se presenta un algoritmo DIRL OPFB sin modelo para aprender iterativamente la función de valor desconocida y la política de control óptima.
- Se realiza un análisis riguroso de la convergencia del algoritmo y la unicidad de la solución.
Principales resultados:
- El algoritmo propuesto recupera eficazmente la política de control experta.
- Las simulaciones demuestran una eficiencia computacional superior en comparación con los métodos del estado del arte.
- El algoritmo maneja con éxito estados parcialmente observables y funciones de valor desconocidas.
Conclusiones:
- El novedoso algoritmo DIRL OPFB proporciona una solución eficaz para controlar sistemas CT desconocidos con información de estado limitada.
- El método mejora la aplicabilidad de DIRL en escenarios prácticos al utilizar solo datos de entrada-salida.
- El algoritmo ofrece un enfoque computacionalmente eficiente y robusto para aprender políticas de control óptimas.
Videos de Conceptos Relacionados
Feedback control systems
703
Feedback control systems are categorized in various ways based on their design, analysis, and signal types.
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
703
Linear time-invariant Systems
890
A system is linear if it displays the characteristics of homogeneity and additivity, together termed the superposition property. This principle is fundamental in all linear systems. Linear time-invariant (LTI) systems include systems with linear elements and constant parameters.
The input-output behavior of an LTI system can be fully defined by its response to an impulsive excitation at its input. Once this impulse response is known, the system's reaction to any other input can be...
The input-output behavior of an LTI system can be fully defined by its response to an impulsive excitation at its input. Once this impulse response is known, the system's reaction to any other input can be...
890
BIBO stability of continuous and discrete -time systems
898
System stability is a fundamental concept in signal processing, often assessed using convolution. For a system to be considered bounded-input bounded-output (BIBO) stable, any bounded input signal must produce a bounded output signal. A bounded input signal is one where the modulus does not exceed a certain constant at any point in time.
To determine the BIBO stability, the convolution integral is utilized when a bounded continuous-time input is applied to a Linear Time-Invariant (LTI) system....
To determine the BIBO stability, the convolution integral is utilized when a bounded continuous-time input is applied to a Linear Time-Invariant (LTI) system....
898
Linear Momentum in Control Volume
1.3K
Newton's second law is applied to obtain the linear momentum in a control volume in a fluid system. According to this law, the rate of change of linear momentum is equal to the sum of external forces acting on the system. When a control volume matches the fluid system at a specific moment, the forces acting on both are identical. Reynolds transport theorem helps explain this by breaking down the system's linear momentum into two components: the rate of change of linear momentum within...
1.3K
Root Loci for Positive-Feedback Systems
338
The Hartley oscillator is a positive feedback system that sustains oscillations by feeding the output back to the input in phase, thereby reinforcing the signal. Positive feedback systems can be viewed as negative feedback systems with inverted feedback signals. In these systems, the root locus encompasses all points on the s-plane where the angle of the system transfer function equals 360 degrees.
The construction rules for the root locus in positive feedback systems are similar to those in...
The construction rules for the root locus in positive feedback systems are similar to those in...
338
Control Systems
1.8K
Control systems are everywhere in contemporary society, influencing diverse applications from aerospace to automated manufacturing. These systems can be found naturally within biological processes, such as blood sugar regulation and heart rate adjustment in response to stress, as well as in man-made systems like elevators and automated vehicles. A control system is essentially a network of subsystems and processes that collaboratively convert specific inputs into desired outputs.
At the heart...
At the heart...
1.8K

