Related Experiment Videos
Less Repetition, Less Energy Cost: A Reinforcement Learning-Based Multiagent Energy-Saving Autonomous Exploration
Abstract:
Multiagent autonomous exploration in unknown environments is both meaningful and challenging. Due to the constraint of a partially observable environment, the collaboration among agents is often inadequate, leading to increased energy consumption. Worse still, a decrease in overall exploration performance may occur due to a single agent failure. To address these issues, we propose a distributed Multiagent Energy-saving Autonomous Exploration System (MEAES) based on reinforcement learning. To accurately evaluate the regional complexity of different branches and further enhance the long-term decision-making capabilities of agents, we introduce the dual-scale clustered observation (DSCO) module. The DSCO generates fine-grained representations based on graph modeling, enabling better characterization of both global and long-term exploration values. Furthermore, we propose an energy-saving action (EA) mechanism, which mitigates redundant exploration and reduces energy consumption by selective waiting actions and independent exploration strategies. Finally, we devise the consumption-exploration-balanced training framework (CEBF), which guides agents to transform from lazy exploration to energy-saving exploration strategies through dynamic reward shaping. Extensive experiments validate the effectiveness of MEAES, demonstrating effective zero-shot transfer performance across unseen environments.
Related Concept Videos
Optimal Foraging
Observational Learning
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Three-Dimensional Force System:Problem Solving
To solve a three-dimensional force system, first resolve each force into its respective scalar components. Do this using...
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Conservation of Mechanical Energy
When a...