相关实验视频
Updated: Jun 24, 2025

04:58
A Rapid Method for Modeling a Variable Cycle Engine
Published on: August 13, 2019
7.6K
揭示了在强风条件下自主热升的原理,使用灵感的深度强化学习
Yoav Flato1,2,3, Roi Harel4,5,6,7, Aviv Tamar8
1Rachel and Selim Benin School of Computer Science and Engineering, The Hebrew University of Jerusalem, Jerusalem, 9190401, Israel.
Nature communications
|June 10, 2024
概括
研究人员使用深度强化学习 (deep-RL) 来训练自主系统进行热升. 该研究确定了学习瓶,并发现神经网络在学习过程中进化了功能集群,为运动控制学习提供了洞察力.
科学领域:
- 机器人技术和自主系统
- 动物行为与生物力学
- 机器学习和人工智能的人工智能
背景情况:
- 使用大气上升的热起,是鸟类观察到的复杂的自然行为,也是自主系统的目标.
- 深度强化学习 (deep-RL) 的最新进展使人工智能能够训练复杂的飞行机动.
- 了解热中的学习过程可以弥合生物和人工运动控制之间的差距.
研究的目的:
- 开发和分析基于模拟的深度RL系统,用于自主热升.
- 调查学习动态,确定瓶,并描述学习政策的稳定性.
- 将新兴的神经网络行为与生物升策略进行比较.
主要方法:
- 开发一个模拟环境用于热升.
- 应用深度强化学习 (deep-RL) 算法来训练自主飞行器.
- 定义和使用一种新的效率指标来评估学习强度.
- 对训练有素的特工政策与飞行的数据进行比较分析.
主要成果:
- 深度RL系统成功学习了自主热升策略.
- 在培训过程中确定了特定的学习瓶.
- 定义了一个新的效率指标,并用于描述学习强度.
- 经过训练的神经网络表现出神经元的功能聚类,这些神经元随着时间的推移而演变,反映了生物控制的各个方面.
结论:
- 热飞行作为一个有效和可处理的模型系统来研究复杂的运动控制的学习.
- 深度强化学习为理解和复制生物飞行策略提供了一个强大的框架.
- 该研究提供了关于神经网络在自主系统中技能获取过程中的内部动态的见解.
相关概念视频
Observational Learning
163
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
163
Neural Control of Respiration
2.4K
The neural regulation of respiration is a meticulously coordinated process primarily controlled by the respiratory centers located within the brainstem. These centers, composed of specialized neurons, transmit nerve impulses that control the contraction and relaxation of our respiratory muscles.
Respiratory Centers in the Brainstem
Two primary areas comprise the respiratory center: the medullary respiratory center in the medulla oblongata and the pontine respiratory group in the pons. The...
Respiratory Centers in the Brainstem
Two primary areas comprise the respiratory center: the medullary respiratory center in the medulla oblongata and the pontine respiratory group in the pons. The...
2.4K
Avoidance Learning and Learned Helplessness
1.7K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.7K
Turbulent Flow: Problem Solving
113
Carbonation is a process used to dissolve carbon dioxide gas in a liquid, commonly used in the production of carbonated beverages. Achieving efficient carbonation requires careful control of temperature, pressure, and flow conditions. By adjusting these parameters, carbonation efficiency can be maximized, producing a higher concentration of CO2 in the liquid.
Temperature is a key factor in CO2 solubility. In this case, the CO2 gas and the liquid are cooled to 20°C. Lower temperatures...
Temperature is a key factor in CO2 solubility. In this case, the CO2 gas and the liquid are cooled to 20°C. Lower temperatures...
113
Reinforcement
202
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
202
Laminar Flow: Problem Solving
159
Laminar flow occurs when a fluid moves smoothly in parallel layers with minimal mixing and turbulence. In fluid mechanics, ensuring laminar flow within a pipe is essential for precise control of flow characteristics, especially in engineering applications. The key factor in determining whether flow remains laminar is the Reynolds number, a dimensionless quantity that depends on the fluid's velocity, density, viscosity, and the pipe's diameter. A Reynolds number of 2100 or lower...
159

