Related Experiment Video
Updated: Aug 29, 2026

Surface Electromyographic Biofeedback as a Rehabilitation Tool for Patients with Global Brachial Plexus Injury Receiving Bionic Reconstruction
Published on: September 28, 2019
Assessment of brachial plexus and upper-extremity peripheral nerve injuries at the anatomical level using multimodal
Vincent G J Guillaume1, Ron Martin2, Jonas Roos3
1Department of Plastic Surgery, Hand and Burn Surgery, University Hospital RWTH Aachen, Aachen, Germany.
Introduction:
Injuries to the brachial plexus and its terminal branches can be devastating at the functional and occupational levels, as loss of even partial functions causes major disruptions in dexterity and mobility. The brachial plexus comprises multiple interconnected anatomical layers, making thorough anatomical knowledge essential for determining the extent of injury and the best treatment. Thus, the correlation between functional deficit and precise localization of the injury site is crucial to avoid extensive surgical explorations and restore anatomical integrity.
Method:
In this study, we presented Medical Research Council (MRC) muscle strength grades from 50 standardized real-world-inspired fictional benchmark cases of brachial plexus and upper-extremity peripheral nerve injuries at different anatomical levels (roots, trunks, cords, terminal branches, and combined patterns) to various Multimodal Generative Models (MM-GMs), namely GPT-5 (OpenAI), Gemini 2.5 Pro (Google), Grok 4 (xAI), and Claude Opus 4.1 (Anthropic), and evaluated their ability to localize anatomical lesion sites based solely on functional motor deficits.
Results:
GPT-5 achieved the highest overall accuracy with 39/50 correct responses (78.0%; 95% CI 64.8-87.2), followed by Grok 4 and Claude Opus 4.1 with 26/50 correct responses each (52.0%; 95% CI 38.5-65.2) and Gemini 2.5 Pro with 22/50 correct responses (44.0%; 95% CI 31.2-57.7). Cochran's Q test showed a significant difference between the models, and Holm-adjusted exact McNemar post-hoc comparisons demonstrated the superior performance of GPT-5 compared with the other models. Combination injuries involving various anatomical sites proved challenging for MM-GMs, with only 20% of answers correct.
Discussion:
Collectively, this controlled benchmark suggests that MM-GMs can assign standardized MRC strength-grade patterns to anatomical levels of brachial plexus and peripheral nerve injuries. However, performance varies substantially across models and lesion categories, requiring validation in real clinical cohorts before clinical implementation.
