Related Experiment Video
Updated: Jul 4, 2026

Synthetic, Multi-Layer, Self-Oscillating Vocal Fold Model Fabrication
Published on: December 2, 2011
Eardrum-inspired soft viscoelastic diaphragms for CNN-based speech recognition with audio visualization images
Seok-Jin Park1, Hee-Beom Lee1, Gi-Woo Kim2
1Department of Mechanical Engineering, Inha University, 100 Inha-ro, Michuhol-gu, Incheon, 22212, Republic of Korea.
Abstract:
In this study, we present initial efforts for a new speech recognition approach aimed at producing different input images for convolutional neural network (CNN)-based speech recognition. We explored the potential of the tympanic membrane (eardrum)-inspired viscoelastic membrane-type diaphragms to deliver audio visualization images using a cross-recurrence plot (CRP). These images were formed by the two phase-shifted vibration responses of viscoelastic diaphragms. We expect this technique to replace the fast Fourier transform (FFT) spectrum currently used for speech recognition. Herein, we report that the new creation method of color images enabled by combining two phase-shifted vibration responses of viscoelastic diaphragms with CRP shows a lower computation burden and a promising potential alternative way to STFT (conventional spectrogram) when the image resolution (pixel size) is below critical resolution.
Related Concept Videos
Hearing
The Cochlea
Perception of Sound Waves
The pitch of a sound depends on the frequency and the pressure amplitude of the source. Two sounds of the same...
Sound as Pressure Waves
The pressure fluctuation depends on the difference in displacements between the successive points in the...
Anatomy of the Ear
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...

