Related Experiment Video
Updated: Sep 26, 2026

Development of a Novel Task-oriented Rehabilitation Program using a Bimanual Exoskeleton Robotic Hand
Published on: May 20, 2020
Biomimetic Dexterous Hand Control for Robotic Piano Playing Using a Two-Stage Reinforcement Learning Curriculum
Lei Jiang1,2, Jinyi Chen1,2, Kaixin Lan1,2
1Center for X-Mechanics, Zhejiang University, Hangzhou 310012, China.
Abstract:
Robotic piano playing is a challenging benchmark for biomimetic dexterous manipulation, requiring precise timing, coordinated multi-finger motion, and stable key contact. This study proposes a robotic piano-playing framework based on a two-stage reinforcement learning curriculum. Musical Instrument Digital Interface (MIDI) data are converted into target-key and fingering grids to provide future musical goals for policy learning in a parallel MJLab simulation environment. A Soft Actor-Critic (SAC) agent takes a 2106-dimensional observation vector, including joint states, previous actions, musical phase, future key targets, fingering assignments, and piano-key states, and outputs a 21-dimensional continuous action vector for wrist, finger, and global hand-positioning control. Stage 1 weakens physical regularization to facilitate key-pressing acquisition, whereas Stage 2 strengthens power, velocity, acceleration, collision, posture, and finger-speed constraints to improve the regularity of policy outputs and readiness for real-world deployment. Simulation experiments on 30 s right-hand excerpts from Für Elise, Canon, and Beethoven's Symphony No. 5 achieve frame-wise key-state F1 scores above 0.99 on the first two excerpts and approximately 0.945 on Beethoven. Real-world deployment uses open-loop playback of policy-generated high-level trajectories with low-level joint-position feedback and achieves F1 scores of 0.95, 0.91, and 0.83, respectively, while reproducing representative piano techniques such as chords, octaves, mixed black-and-white-key patterns, overlapping finger actions, and rapid sequential movements. These physical results demonstrate the feasibility of the proposed sim-to-real pipeline for complete 30 s executions; they are not intended as a statistical repeatability study. The results further show that biomimetic robotic hands can learn complex piano-playing skills from MIDI-based task objectives without relying on human motion demonstration trajectories.

