Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Hearing01:31

Hearing

52.6K
When we hear a sound, our nervous system is detecting sound waves—pressure waves of mechanical energy traveling through a medium. The frequency of the wave is perceived as pitch, while the amplitude is perceived as loudness.
52.6K
The Cochlea01:13

The Cochlea

45.4K
The cochlea is a coiled structure in the inner ear that contains hair cells—the sensory receptors of the auditory system. Sound waves are transmitted to the cochlea by small bones attached to the eardrum called the ossicles, which vibrate the oval window that leads to the inner ear. This causes fluid in the chambers of the cochlea to move, vibrating the basilar membrane.
45.4K
Perception of Sound Waves01:01

Perception of Sound Waves

4.5K
The human ear is not equally sensitive to all frequencies in the audible range. It may perceive sound waves with the same pressure but different frequencies as having different loudness. Moreover, the perception of sound waves depends on the health of an individual's ears, which decays with age. The health of one's ears may also be affected by regular exposure to loud noises.
The pitch of a sound depends on the frequency and the pressure amplitude of the source. Two sounds of the same...
4.5K
Sound as Pressure Waves01:17

Sound as Pressure Waves

2.5K
Sound waves, which are longitudinal waves, can be modeled as the displacement amplitude varying as a function of the spatial and temporal coordinates. As a column of the medium is displaced, its successive columns are also displaced. As the successive displacements differ relatively, a pressure difference with the surrounding pressure is created. The gauge pressure varies across the medium.
The pressure fluctuation depends on the difference in displacements between the successive points in the...
2.5K
Anatomy of the Ear01:16

Anatomy of the Ear

8.6K
Auditory sensation, commonly called hearing, involves the transformation of sonic waves into neural impulses facilitated by the structures of the auditory organ. The prominent, flesh-like structure on the side of the head, called the auricle, directs sound waves towards the auditory canal. The auricle is often mislabeled as the pinna, a term more aligned with mobile structures like a feline's external ear. The auditory canal penetrates the cranium via the external auditory meatus of the...
8.6K
Perceiving Loudness, Pitch, and Location01:21

Perceiving Loudness, Pitch, and Location

287
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
287

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Homoepitaxy-like heteroepitaxy via monolayer interface achieves grain-boundary-free ultraflat silver thin films.

Reports on progress in physics. Physical Society (Great Britain)·2026
Same author

A Trailblazing Quenching Strategy for Simultaneous LiF Formation at Surface and Intergranular Interfaces for Enhanced Stability of High-Ni NCM Cathodes.

Small (Weinheim an der Bergstrasse, Germany)·2025
Same author

FIB-SEM: Emerging Multimodal/Multiscale Characterization Techniques for Advanced Battery Development.

Chemical reviews·2025
Same author

Crack Monitoring in Rotating Shaft Using Rotational Speed Sensor-Based Torsional Stiffness Estimation with Adaptive Extended Kalman Filters.

Sensors (Basel, Switzerland)·2023
Same author

Tunable Young's Moduli of Soft Composites Fabricated from Magnetorheological Materials Containing Microsized Iron Particles.

Materials (Basel, Switzerland)·2020
Same author

Microstructure Simulation and Constitutive Modelling of Magnetorheological Fluids Based on the Hexagonal Close-packed Structure.

Materials (Basel, Switzerland)·2020

Related Experiment Video

Updated: Jul 4, 2026

Synthetic, Multi-Layer, Self-Oscillating Vocal Fold Model Fabrication
10:16

Synthetic, Multi-Layer, Self-Oscillating Vocal Fold Model Fabrication

Published on: December 2, 2011

Eardrum-inspired soft viscoelastic diaphragms for CNN-based speech recognition with audio visualization images.

Seok-Jin Park1, Hee-Beom Lee1, Gi-Woo Kim2

  • 1Department of Mechanical Engineering, Inha University, 100 Inha-ro, Michuhol-gu, Incheon, 22212, Republic of Korea.

Scientific Reports
|April 19, 2023
PubMed
Summary

Researchers developed a new way to recognize speech by mimicking the human eardrum. They used flexible, vibrating membranes to turn sound into visual patterns, which are then analyzed by artificial intelligence. This method could eventually replace standard audio processing techniques by being faster and more efficient for certain image sizes.

Keywords:
bio-inspired sensorsconvolutional neural networksacoustic signal processingcross-recurrence plots

Frequently Asked Questions

More Related Videos

Construction and Characterization of a Novel Vocal Fold Bioreactor
11:11

Construction and Characterization of a Novel Vocal Fold Bioreactor

Published on: August 1, 2014

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
05:48

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis

Published on: August 9, 2024

Related Experiment Videos

Last Updated: Jul 4, 2026

Synthetic, Multi-Layer, Self-Oscillating Vocal Fold Model Fabrication
10:16

Synthetic, Multi-Layer, Self-Oscillating Vocal Fold Model Fabrication

Published on: December 2, 2011

Construction and Characterization of a Novel Vocal Fold Bioreactor
11:11

Construction and Characterization of a Novel Vocal Fold Bioreactor

Published on: August 1, 2014

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
05:48

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis

Published on: August 9, 2024

Area of Science:

  • Acoustics and signal processing within viscoelastic diaphragms research
  • Machine learning applications in speech recognition systems

Background:

No prior work had resolved how to effectively utilize biological inspiration to simplify complex audio processing tasks for machine learning. It was already known that standard spectral analysis requires significant computational power for real-time speech recognition. That uncertainty drove researchers to look for alternatives to traditional mathematical transformations of sound waves. Prior research has shown that the human auditory system processes sound through mechanical vibrations of the tympanic membrane. This gap motivated the development of synthetic structures that mimic these natural physical responses. Scientists have long sought ways to reduce the processing load required for convolutional neural networks to interpret human speech. No previous studies had combined viscoelastic material properties with specific visual plotting techniques for this purpose. This investigation addresses the need for more efficient input data formats for modern artificial intelligence models.

Purpose Of The Study:

The aim of this study is to introduce a novel speech recognition approach that utilizes bio-inspired diaphragms to generate input images for neural networks. Researchers sought to address the high computational costs associated with standard audio processing techniques. They specifically aimed to replace the fast Fourier transform spectrum with a more efficient, membrane-based method. The team investigated whether mimicking the human eardrum could simplify the conversion of sound into machine-readable data. They hypothesized that viscoelastic properties would provide unique vibration responses suitable for visual analysis. This work was motivated by the need for faster, less resource-intensive speech recognition systems in artificial intelligence. The authors intended to demonstrate that cross-recurrence plots could effectively represent audio data for convolutional neural networks. This research addresses the gap in developing hardware-level solutions for optimizing speech data input formats.

Main Methods:

The review approach involved designing synthetic membranes that replicate the mechanical behavior of the human tympanic membrane. Investigators utilized viscoelastic materials to ensure the diaphragms could produce distinct vibration responses. The team implemented a system to capture two phase-shifted signals from these vibrating structures. These signals were then transformed into visual representations using cross-recurrence plotting techniques. The researchers evaluated the resulting images as inputs for convolutional neural network architectures. They compared the computational requirements of this new method against traditional spectral analysis tools. The study focused on testing performance across various image resolutions to identify operational limits. This systematic evaluation provided data on the efficiency of the bio-inspired sensing approach.

Main Results:

The strongest finding indicates that the proposed method significantly lowers the computational burden compared to traditional spectral analysis. The researchers observed that combining two phase-shifted responses with cross-recurrence plots creates effective input images for neural networks. This technique serves as a promising alternative to conventional spectrograms when image resolution stays below the critical threshold. The data show that the mechanical properties of the diaphragms successfully translate sound into visual patterns. These patterns allow convolutional neural networks to perform speech recognition tasks without standard Fourier-based processing. The authors report that the system maintains functionality while reducing the complexity of data preparation. The findings suggest that the pixel size of the generated images directly influences the efficiency gains observed. This work demonstrates that bio-inspired hardware can successfully interface with modern machine learning models.

Conclusions:

The authors suggest that their membrane-based approach offers a viable alternative to standard spectral analysis for speech recognition. They propose that combining phase-shifted responses with visual plotting reduces the overall computational burden. The team claims this method performs effectively when image resolution remains below a specific threshold. This synthesis indicates that bio-inspired hardware can simplify data preparation for machine learning tasks. The researchers conclude that their technique provides a unique way to generate input images for neural networks. They highlight that this approach avoids the heavy processing requirements of traditional Fourier-based methods. The study implies that viscoelastic properties are beneficial for capturing sound characteristics in a format suitable for visual analysis. These findings support the potential integration of mechanical sensors into future speech recognition architectures.

The researchers propose using viscoelastic diaphragms to capture two phase-shifted vibration responses. These signals are converted into cross-recurrence plots, which serve as input images for convolutional neural networks, offering a lower computational load compared to traditional fast Fourier transform methods.

The study utilizes cross-recurrence plots, a technique that visualizes the relationship between two time-series signals. Unlike standard spectrograms, this method maps the phase-shifted mechanical vibrations of the synthetic eardrum to create distinct patterns for machine learning classification.

The authors state that this specific resolution is necessary because the proposed method shows superior performance compared to conventional spectrograms only when the pixel size falls below this critical limit, highlighting a boundary for the technique's efficiency.

These diaphragms act as the primary sensing component, mimicking the human eardrum to convert sound waves into physical vibrations. Their viscoelastic nature allows for the generation of the specific phase-shifted responses required to build the visual data inputs.

The researchers measure the vibration responses of the synthetic membranes to generate color images. This phenomenon relies on the interaction between the two phase-shifted signals, which are then processed to evaluate the system's potential as an alternative to standard short-time Fourier transform spectrograms.

The authors propose that this bio-inspired hardware could eventually replace the fast Fourier transform spectrum currently used in speech recognition. They suggest that this shift would provide a more efficient, lower-burden pathway for processing audio data in artificial intelligence applications.