Related Experiment Video
Updated: Jun 9, 2026

07:42
An Automated T-maze Based Apparatus and Protocol for Analyzing Delay- and Effort-based Decision Making in Free Moving Rodents
Published on: August 2, 2018
Data-driven simulator of multi-animal behavior with unknown dynamics via reinforcement learning.
Keisuke Fujii1,2, Kazushi Tsutsui3,2, Yu Teshima4
1Nagoya University, Nagoya, Japan.
Iscience
|June 8, 2026
Summary
This study introduces a novel data-driven simulator for multi-animal behavior using deep reinforcement learning. It achieves higher reproducibility and enables counterfactual predictions for complex biological systems.
Area of Science:
- Robotics and Artificial Intelligence
- Computational Biology
- Behavioral Ecology
Background:
- Advances in imitation learning enable human and animal movement replication.
- Simulating realistic multi-animal behaviors is challenging due to unknown transition models and locomotion dynamics.
- Existing methods struggle to construct simulators that reproduce trajectories and support reward-driven optimization.
Purpose of the Study:
- To introduce a data-driven simulator for multi-animal behavior.
- To address challenges in simulating complex biological systems with unknown dynamics.
- To enable reward-driven optimization and counterfactual behavior prediction.
Main Methods:
- Utilized deep reinforcement learning with counterfactual simulations for a data-driven approach.
- Estimated movement variables in reinforcement learning to handle high degrees of freedom.
- Employed a distance-based pseudo-reward to align cyber and physical states.
Main Results:
- Achieved higher reproducibility and reward acquisition compared to imitation learning and reinforcement learning.
- Validated the approach using data from artificial agents, flies, newts, and silkmoths.
- Demonstrated the capability for counterfactual behavior prediction in novel experimental settings.
Conclusions:
- The developed simulator shows significant potential for understanding complex multi-animal behaviors.
- This data-driven approach overcomes limitations of traditional mathematical models in biological simulations.
- Enables advanced analysis and prediction of animal interactions and collective dynamics.
Related Concept Videos
Reinforcement Schedules
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
Observational Learning
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning because...
Multi-input and Multi-variable systems
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence of...
In the absence of...

