Related Experiment Video
Updated: Mar 29, 2026

Multi-Modal Signals for Analyzing Pain Responses to Thermal and Electrical Stimuli
Published on: April 5, 2019
AVPENet: Pain estimation from audio-visual fusion of non-speech sounds
Sami Naouali1, Oussama El Othmani2,3
1Information Systems Department, College of Computer Science and Information Technology, King Faisal University, Al Ahsa, Saudi Arabia.
Abstract:
Pain assessment in non-verbal patients, including neonates and unconscious adults, remains a critical challenge in clinical practice. Current pain scales rely heavily on observer interpretation and may lack objectivity, introducing significant inter-rater variability. We propose a novel multimodal deep learning framework that estimates continuous pain intensity by fusing non-speech audio cues with facial expressions. Our approach addresses the critical need for objective pain assessment in vulnerable populations unable to self-report. We developed a cross-modal attention-based fusion network combining spectrogram-derived audio embeddings with facial action unit features. The model was trained and validated on 3,247 audio-visual recordings from 428 subjects, including 215 neonates and 213 adults, across three distinct pain intensity levels. We employed a ResNet-based audio encoder for mel-spectrogram processing and a facial landmark convolutional neural network for expression analysis, integrated through a transformer-based fusion module that learns complementary relationships between modalities. Our model achieved a mean absolute error of 0.89 on a 0-10 pain scale, significantly outperforming audio-only approaches (mean absolute error 1.47, 39% improvement) and visual-only baselines (mean absolute error 1.23, 28% improvement). Cross-age group validation demonstrated robust generalization with mean absolute errors of 0.94 for neonates and 0.91 for adults. The model maintained a Pearson correlation coefficient of 0.89 with ground truth annotations and achieved 81.4% accuracy for three-class pain categorization. Audio-visual fusion significantly enhances pain estimation accuracy across diverse age groups and clinical scenarios. This approach offers substantial potential for objective, automated pain monitoring in clinical settings, particularly for vulnerable populations unable to self-report pain.
More Related Videos
07:28Psychophysically-anchored, Robust Thresholding in Studying Pain-related Lateralization of Oscillatory Prestimulus Activity
Published on: January 21, 2017
03:58Enhancing Electrode Location Assessment in Cochlear Implantation via Computed Tomography Image Fusion
Published on: January 17, 2025
Related Concept Videos
Sound Intensity Level
The human ear can perceive an extensive range of sound intensity, necessitating the use of the logarithmic scale to define a physical quantity—the intensity level. It is a ratio of two intensities and...
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Perception of Sound Waves
The pitch of a sound depends on the frequency and the pressure amplitude of the source. Two sounds of the same...
Auditory Perception
Hearing
Sound Intensity