Related Experiment Video
Updated: Aug 21, 2026

Estimation of Contact Regions Between Hands and Objects During Human Multi-Digit Grasping
Published on: April 21, 2023
Assessment of open surgery suturing skills via markerless hand tracking: a multistage approach using 3D hand skeleton
Noé Constans1,2, Mihai Mitrea3, Nicole Neumann4
1Télécom SudParis, Institut Polytechnique de Paris, Palaiseau, France. noe.constans@ip-paris.fr.
Purpose:
The transition toward competency-based medical education requires scalable, objective surgical skill assessment. While sensor-based approaches remain burdensome, video-based deep learning models often lack the interpretability required for formative feedback. We introduce a fully automated, markerless framework for assessing vascular open surgery suturing skills providing meaningful interpretation.
Methods:
Our pipeline leverages 3D hand tracking coupled to an LSTM-attention network for precise temporal segmentation. By merging explicit kinematic metrics with latent deep features, we develop a voting ensemble classifier categorizing surgeons into three skill levels. We further evaluated the impact of ground-truth reliability by comparing models trained on single-assessor versus multi-assessor consensus labels.
Results:
The framework achieved 96.4% segmentation accuracy and 72.2 ± 0.5% accuracy (F1 = 0.73) under leave-one-video-out cross-validation (LOVCV). To go deeper in assessing cross-surgeon generalization, we additionally report a strict leave-one-subject-out (LOSO) protocol, under which accuracy decreases to 52.3%. This gap quantifies the surgeon-specific component captured by video-level splits and reframes the system as a tool for longitudinal, per-trainee skill tracking rather than zero-shot scoring of unseen surgeons. Phase-stratified analysis uncovers that experts exhibit significantly higher jerk and velocity, reflecting decisive ballistic motor planning. Training on consensus labels yields a 15-20% performance increase over single-assessor baselines.
Conclusion:
Using a single camera and no physical markers, this framework offers a cost-effective alternative to hardware-heavy simulators, enabling democratized, interpretable feedback.
