Related Experiment Video
Updated: Sep 10, 2025

Automated Interactive Video Playback for Studies of Animal Communication
Published on: February 9, 2011
Maynard Smith revisited: A multi-agent reinforcement learning approach to the coevolution of signalling behaviour
Olivia Macmillan-Scott1, Mirco Musolesi1,2
1AI Centre, Department of Computer Science, University College London, London, United Kingdom.
None:
The coevolution of signalling is a complex problem within animal behaviour, and is also central to communication between artificial agents. The Sir Philip Sidney game was designed to model this dyadic interaction from an evolutionary biology perspective, and was formulated to demonstrate the emergence of honest signalling. We use Multi-Agent Reinforcement Learning (MARL) to show that in the majority of cases, the resulting behaviour adopted by agents is not that shown in the original derivation of the model. This paper demonstrates that MARL can be a powerful tool to study evolutionary dynamics and understand the underlying mechanisms of learning over generations; particularly advantageous is the interpretability of this type of approach, as well as the fact that it allows us to study emergent behaviour without the need to constrain the strategy space from the outset. Although it originally set out to exemplify honest signalling, we show that the game provides no incentive for such behaviour. In the majority of cases, the optimal outcome is one that does not require a signal for the resource to be given. This type of interaction is observed within animal behaviour and is sometimes referred to as proactive prosociality. High learning and low discount rates of the reinforcement learning model are shown to be optimal in order to achieve the outcome that maximises both agents' reward, and proximity to the given threshold leads to suboptimal learning.
Related Concept Videos
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Observational Learning
Interactions Between Signaling Pathways
Convergence and divergence, and cross-talk between signaling pathways
Two distinct signaling pathways can converge on a single functional unit, which may either be a single protein or a complex of proteins. The response is either functionally distinct or synergistic between the two pathways but different from the response...
Reinforcement Schedules
Once a behavior is learned,...
Autocrine Signaling
Autocrine Signaling in Macrophages
Under normal physiological conditions, autocrine signaling is essential for maintaining homeostasis. This process is well characterized in...
Evolutionary Psychology

