使用强化学习的平流层气球的自主导航
Marc G Bellemare1, Salvatore Candido2, Pablo Samuel Castro3
1Brain Team, Google Research, Montreal, Quebec, Canada. bellemare@google.com.
Nature
|December 3, 2020
概括
强化学习可以实现超压气球的自主飞行控制,克服数据的缺陷. 这种由人工智能驱动的系统比以前的方法更有效地导航平流层.
科学领域:
- 人工智能
- 航空航天工程
- 自主系统
背景情况:
- 气球导航面临风力变化,预测错误和稀缺数据等挑战.
- 传统的控制方法不足以在动态的平流层环境中进行实时决策.
研究的目的:
- 通过强化学习开发高性能飞行控制器用于超压气球.
- 解决物理系统中不完美的数据强化学习的挑战.
主要方法:
- 使用强化学习 (RL) 进行数据增强和自我纠正设计.
- 在全球范围内部署RL控制器.
- 在太平洋上进行了为期39天的受控实验.
主要成果:
- 与Loon之前的算法相比,RL控制器表现出更高的性能.
- 控制器在不同平流层风的条件下被证明是坚固的.
- 成功地克服了将RL应用到不完美的现实数据的障碍.
结论:
- 强化学习在复杂的现实场景中提供了一种有效的自主控制解决方案.
- 这种方法适用于需要与动态环境持续交互的系统.
- 在自主导航和控制方面为先进的人工智能机构铺平道路.
更多相关视频
08:37A Video Demonstration of Preserved Piloting by Scent Tracking but Impaired Dead Reckoning After Fimbria-Fornix Lesions in the Rat
Published on: April 24, 2009
12.1K
07:17Author Spotlight: A Novel Standardized Technique for Real-Time Biomedical Imaging of Acute Myocardial Injury
Published on: March 22, 2024
1.4K
相关概念视频
Reinforcement
633
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
633
Buoyancy and Stability for Submerged and Floating Bodies
2.3K
In fluid mechanics, buoyancy and stability are key concepts for understanding the behavior of submerged and floating bodies. When a stationary body is fully or partially submerged in a fluid, the fluid exerts a force on the body known as the buoyant force. This force acts vertically upward through a point called the center of buoyancy, which is the center of the displaced fluid volume. According to Archimedes' principle, the magnitude of the buoyant force is equal to the weight of the fluid...
2.3K
Neural Control of Respiration
3.9K
The neural regulation of respiration is a meticulously coordinated process primarily controlled by the respiratory centers located within the brainstem. These centers, composed of specialized neurons, transmit nerve impulses that control the contraction and relaxation of our respiratory muscles.
Respiratory Centers in the Brainstem
Two primary areas comprise the respiratory center: the medullary respiratory center in the medulla oblongata and the pontine respiratory group in the pons. The...
Respiratory Centers in the Brainstem
Two primary areas comprise the respiratory center: the medullary respiratory center in the medulla oblongata and the pontine respiratory group in the pons. The...
3.9K
Observational Learning
663
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
663
Reinforcement Schedules
341
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
341
Flat Belts: Problem Solving
631
Flat belts are crucial in many industrial applications as they help transmit power from one pulley to another. The concept of forces and moments is used to determine the maximum moment on a pulley. For instance, consider a flat belt that wraps around two pulleys, A and B, with radii of 30 cm and 10 cm, respectively. The angle between the belt and the horizontal is 20 degrees at the pulleys. As pulley B rotates clockwise and drives pulley A, tension T2 is caused at one end of the belt, while...
631
