Related Experiment Video
Updated: Aug 6, 2026

Multi-Modal Signals for Analyzing Pain Responses to Thermal and Electrical Stimuli
Published on: April 5, 2019
Deciphering the "non-verbal" code: A preliminary exploration of multimodal large language models for neonatal pain
Xiaosong Jiang1, Nannan Yang2, Tingqi Shi3
1Department of Intensive Care Medicine, The First Affiliated Hospital of Soochow University, Suzhou, China.
Background:
Neonatal pain assessment primarily relies on behavioral rating scales. Although these tools provide objective measures, their scores are susceptible to inter-rater variability, potentially introducing bias. Moreover, intermittent assessments cannot capture pain progression over time. Multimodal large language models (MLLMs) offer a novel approach to addressing these limitations; however, their effectiveness in neonatal pain recognition remains largely unexplored.
Objective:
This study aimed to evaluate the performance of several MLLMs in neonatal pain video classification and investigate their applicability, limitations, and potential for clinical implementation.
Methods:
A previously established, annotation-validated multimodal dataset of acute neonatal pain was used. The dataset comprised 426 video recordings of neonatal heel lance procedures. Three MLLMs with dynamic video analysis capabilities (Qwen3-VL-Plus, Gemini-3-Pro, and ERNIE-4.5-Turbo) were evaluated. Carefully designed prompts guided the models to simulate nurses' pain assessments. Model performance was comprehensively evaluated using accuracy, weighted Cohen's kappa coefficient, precision, recall, and F1 score.
Results:
Gemini-3-Pro achieved the best classification performance, with an accuracy of 86.9% and a weighted kappa coefficient of 0.695, followed by Qwen3-VL-Plus (83.6%, kappa = 0.624) and ERNIE-4.5-Turbo (78.4%, kappa = 0.563). Gemini-3-Pro demonstrated an excellent recall of 0.974 and an F1 score of 0.905 for the pain category. However, all models showed relatively lower recall for the no-pain category (0.664-0.868), suggesting a trade-off between pain and no-pain classification and indicating challenges in achieving both high sensitivity and specificity.
Conclusion:
MLLMs demonstrate considerable potential for neonatal pain assessment, with particularly strong performance in pain screening. Their structured reasoning and interpretable outputs may support clinical decision-making. However, the observed performance imbalance between pain and no-pain classification suggests that these models are not yet capable of independently replacing clinical judgment and are better suited as assistive tools to optimize clinical workflows.

