Related Experiment Video
Updated: Dec 20, 2025

Real-Time Proxy-Control of Re-Parameterized Peripheral Signals using a Close-Loop Interface
Published on: May 8, 2021
Speech Discrimination in Real-World Group Communication Using Audio-Motion Multimodal Sensing
Takayuki Nozawa1, Mizuki Uchiyama2, Keigo Honda2
1Research Institute for the Earth Inclusive Sensing, Tokyo Institute of Technology, Tokyo 152-8550, Japan.
Combining audio and physical motion data from wearable smartphones significantly improves speech discrimination in group settings. This multimodal sensing approach enhances the ability to identify individual utterances amidst simultaneous speech, crucial for understanding human communication.
Area of Science:
- Human-Computer Interaction
- Signal Processing
- Communication Systems
Background:
- Speech discrimination is vital for analyzing human verbal communication.
- Simultaneous speakers in dynamic group settings pose challenges for audio-only speech detection.
- Existing methods struggle to accurately distinguish individual speech in multi-participant environments.
Purpose of the Study:
- To investigate if combining audio and physical motion data improves speech discrimination.
- To assess the effectiveness of multimodal sensing using wearable smartphones for utterance detection.
- To enhance speech discrimination in dynamic group communication scenarios.
Main Methods:
- Recorded utterance and physical activity data from university students using neck-worn smartphones.
- Analyzed the temporal relationship between physical activities and speech utterances.
- Trained and tested audio-only, motion-only, and combined audio-motion classifiers for speech discrimination.
Main Results:
- Physical activities across a wide frequency range were found to co-occur with speech utterances.
- The audio-motion classifier achieved higher accuracy (92.2%) compared to audio-only (80.4%) and motion-only (87.8%) for intra-participant classification.
- The combined classifier also outperformed single-modality classifiers in inter-individual classification (83.2% vs. 67.7% and 71.9%).
Conclusions:
- Multimodal sensing combining audio and motion data effectively improves speech discrimination.
- Widely available smartphones can be utilized for robust utterance detection in dynamic group communications.
- This approach offers a promising solution for analyzing complex social interactions and verbal communication.
More Related Videos
05:48Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
11:39Assessment of Audio-Tactile Sensory Substitution Training in Participants with Profound Deafness Using the Event-Related Potential Technique
Published on: September 7, 2022
Related Concept Videos
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Perception of Sound Waves
The pitch of a sound depends on the frequency and the pressure amplitude of the source. Two sounds of the same...
Sensory Modalities
General senses refer to the broad category of sensory information detected by receptors in the body and can be further grouped into somatic and visceral senses. Somatic sensations include touch, pressure, temperature, and pain and are essential for navigating our environment and...
Auditory Perception