Related Experiment Video
Updated: Oct 11, 2025

Robotic Mirror Therapy System for Functional Recovery of Hemiplegic Arms
Published on: August 15, 2016
Imitation and mirror systems in robots through Deep Modality Blending Networks
M Yunus Seker1, Alper Ahmetoglu1, Yukie Nagai2
1Bogazici University, Bebek, Istanbul, 34342, Turkey.
This study introduces deep modality blending networks (DMBN) for robots to understand actions and imitate behaviors by creating a shared latent space from multi-modal experiences. DMBN enables robust mirror learning, outperforming vision-only models, by integrating proprioceptive and visual data.
Area of Science:
- Robotics and Artificial Intelligence
- Machine Learning
- Computational Neuroscience
Background:
- Robots learning to interact with environments enhances manipulation, action understanding, and imitation.
- Biological systems, like primates with mirror neurons, demonstrate multi-modal action understanding.
- Enabling robots to leverage interaction experience for understanding others' actions remains a challenge.
Purpose of the Study:
- To propose a novel method, deep modality blending networks (DMBN), for creating a common latent space from multi-modal robot experiences.
- To demonstrate DMBN's capability in facilitating action recognition and enabling anatomical and effect-based imitation.
- To establish a computational model for mirror neuron-like capabilities and a robust machine learning architecture for multi-modal temporal data.
Main Methods:
- Developed deep modality blending networks (DMBN) using a stochastic weighting mechanism to blend multi-modal signals into a common latent space.
- Utilized conditional neural processes allowing conditioning on any sensory/motor value for one-shot generation of complete multi-modal trajectories.
- Conducted simulation experiments with an arm-gripper robot and RGB camera, comparing DMBN with multi-modal variational autoencoders.
Main Results:
- DMBN accurately predicted missing modalities (camera or joint angles) and outperformed multi-modal variational autoencoders in long-horizon trajectory predictions.
- The system generated corresponding image and joint angle sequences for anatomical or effect-based imitation based on desired images.
- Mirror learning was achieved in all DMBN scenarios, whereas vision-only models failed in half, highlighting the necessity of proprioceptive experience.
Conclusions:
- DMBN provides a novel approach to action recognition and imitation by creating a shared latent space from multi-modal robot interaction data.
- The proposed architecture effectively enables mirror neuron-like behavior without pixel-based matching, relying on the blended latent space.
- DMBN serves as a powerful machine learning architecture for high-dimensional, multi-modal temporal data, demonstrating robust retrieval with partial information.
Related Concept Videos
Nonconscious Mimicry
Modeling and Similitude
Observational Learning
Multi-input and Multi-variable systems
In the absence...
Stereotype Content Model
Automatic Processing and Automatic Social Behavior

