Related Experiment Video
Updated: Sep 26, 2026

Utilizing vmTracking to Improve the Accuracy of Multi-Animal Pose Estimation in Rodent Social Behavior Studies
Published on: November 7, 2025
Multi-Camera Analysis of Simulated Porcine Recovery States Based on Three-Dimensional Pose Using Synthetic Data
Artem Obukhov1,2, Daniil Teselkin1, Maxim Shiltsyn1
1Laboratory of VR Simulators, Tambov State Technical University, Tambov 392000, Russia.
Abstract:
Continuous objective assessment of pig motor condition, including recovery after anesthesia, requires the joint analysis of posture, locomotion, transitions between states, and short-term adverse events. This study aimed to develop and algorithmically evaluate a multi-camera pipeline for the quantitative analysis of simulated recovery states in a fully synthetic virtual environment. Procedural generation was used to produce 150,000 images annotated with fifteen anatomical keypoints while varying the pose, size, and appearance of the model, camera viewpoints, illumination, environment, and image post-processing parameters. The YOLO11m-pose model was used for two-dimensional pose estimation, after which observations from three synchronized cameras were combined using weighted triangulation. The reconstructed three-dimensional trajectories were processed using a quality-control system, a finite-state machine, and temporal rules for detecting falls, prolonged immobility, and convulsion-like movements. After repartitioning the dataset by three-dimensional pose index, the retrained YOLO11m-pose model achieved a keypoint mAP50-95 of 0.8968, a precision of 0.9985, and a recall of 0.9986 on the independent pose-level test set. During the processing of a 15-min three-camera sequence, 98.993% of 405,000 reconstructions satisfied the geometric acceptance criteria, and the median reprojection error was 2.420 pixels. On a first manually annotated 5-min synthetic sequence, the state machine achieved a strict frame-level accuracy of 91.20%, a macro-F1 score of 90.10%, and Cohen's κ of 0.860; accuracy outside ±1 s neighborhoods of state transitions was 97.43%. On a second 5-min synthetic evaluation sequence with exact Unity-world geometric ground truth, direct three-dimensional evaluation yielded an MPJPE of 31.86 mm at 99.993% landmark coverage. A limited temporal assessment on this second sequence detected both of the two lying-derived prolonged-immobility reference intervals; fall and convulsion-like-movement events were not independently annotated. All training, validation, and end-to-end evaluation data were generated in a virtual environment; therefore, these results establish algorithmic feasibility within the synthetic domain but do not establish performance on real animals.