Related Experiment Video
Updated: Aug 5, 2026

Insect-machine Hybrid System: Remote Radio Control of a Freely Flying Beetle (Mercynorrhina torquata)
Published on: September 2, 2016
A Comparative Study of Control Approaches in Hybrid Reinforcement Learning-Based Drone Swarms
Raúl Arranz1, Juan A Besada1, David Carramiñana1
1Information Processing and Telecommunications Center, Universidad Politécnica de Madrid, ETSI Telecomunicación, Av. Complutense 30, 28040 Madrid, Spain.
Abstract:
Reinforcement learning (RL) has emerged as a powerful paradigm for enabling autonomous coordination in multi-UAV systems operating in complex and uncertain environments. However, the effectiveness of learned policies is strongly influenced by how actions are implemented at the control level, an aspect that has received limited attention in the literature. This paper presents a comparative study of three control methods (heading-based, waypoint-based, and deterministic) within a unified hybrid-AI architecture, in which the same RL policy structure is used across two of the three configurations. By isolating the control method as the sole variable, the study evaluates how different action abstractions affect learning efficiency, robustness, and operational performance in cooperative surveillance missions. A statistically rigorous Monte Carlo evaluation, supported by non-parametric hypothesis testing, demonstrates that heading-based control consistently achieves superior performance in terms of revisit period, target acquisition time, and tracking continuity. The analysis further reveals that these gains arise from improved reactivity and constraint handling rather than from differences in policy learning. The results highlight the critical role of control-level design in RL-based multi-agent systems and provide practical guidelines for selecting action abstractions in aerial swarm applications.
