Related Experiment Video
Updated: May 8, 2025

10:13
Hemi-laryngeal Setup for Studying Vocal Fold Vibration in Three Dimensions
Published on: November 25, 2017
10.9K
Laryngeal disease classification using voice data: Octave-band vs. mel-frequency filters
Jaemin Song1, Hyunbum Kim2, Yong Oh Lee1
1Department of Industrial and Data Engineering, Hongik University, Seoul, South Korea.
Heliyon
|December 25, 2024
Summary
Octave Frequency Spectrum Energy (OFSE) with 1/3 octave band filters shows higher accuracy than Mel Frequency Cepstral Coefficients (MFCCs) for diagnosing laryngeal diseases. This AI-driven voice analysis offers a promising non-invasive method for early laryngeal cancer detection.
Area of Science:
- Otolaryngology
- Biomedical Engineering
- Artificial Intelligence in Healthcare
Background:
- Laryngeal cancer diagnosis traditionally relies on specialist examinations.
- Non-invasive diagnostic methods using voice analysis are emerging with AI advancements.
- Mel Frequency Cepstral Coefficients (MFCCs) are standard for voice analysis, but may lack accuracy for subtle laryngeal changes.
Purpose of the Study:
- To compare the diagnostic effectiveness of MFCCs and Octave Frequency Spectrum Energy (OFSE) in classifying voice data.
- To differentiate between healthy voices, laryngeal cancer, benign mucosal disease, and vocal fold paralysis using AI.
- To evaluate OFSE with 1/3 octave band filters for improved laryngeal disease detection.
Main Methods:
- Analysis of voice samples from 363 patients using Convolutional Neural Network (CNN) models.
- Application of both MFCC and OFSE with 1/3 octave band filters for feature extraction.
- Utilized Grad-Class Activation Mapping (Grad-CAM) for visualizing critical voice features.
Main Results:
- OFSE with 1/3 octave band filters achieved significantly higher classification accuracy (0.9398 ± 0.0232) compared to MFCCs (0.7061 ± 0.0561).
- Grad-CAM identified distinct voice features for laryngeal cancer, including noise in the over-formant area and fundamental frequency changes.
- AI differentiation between benign conditions and laryngeal cancer remains challenging due to overlapping voice characteristics.
Conclusions:
- OFSE with 1/3 octave band filters demonstrates superior performance for diagnosing laryngeal diseases, including cancer.
- This AI-based voice analysis method shows potential for accurate, non-invasive early detection of laryngeal conditions.
- Further research is needed to address challenges in differentiating benign mucosal diseases from laryngeal cancer using voice data.
Related Concept Videos
Larynx
1.1K
The human larynx, often referred to as the voice box, is an intricate organ located in the neck. It serves as a pathway for air to enter the lungs during respiration and is an essential component of voice production.
Anatomy of the Larynx
The larynx consists of various components, including cartilage, muscles, and vocal cords. Its structure includes three large unpaired cartilages—the thyroid, cricoid, and epiglottis—and three smaller paired cartilages—the arytenoids,...
Anatomy of the Larynx
The larynx consists of various components, including cartilage, muscles, and vocal cords. Its structure includes three large unpaired cartilages—the thyroid, cricoid, and epiglottis—and three smaller paired cartilages—the arytenoids,...
1.1K
Passive Filters
420
Passive filters are utilized to shape the frequency spectrum of signals across a diverse array of applications. These filters, using only passive elements like resistors (R), inductors (L), and capacitors (C), are capable of selectively allowing or blocking certain frequency ranges without the need for external power sources.
Low-Pass Filters
Low-pass filters are designed to transmit signals with frequencies lower than the cutoff frequency, ωc, and attenuate those above it. The cutoff...
Low-Pass Filters
Low-pass filters are designed to transmit signals with frequencies lower than the cutoff frequency, ωc, and attenuate those above it. The cutoff...
420
Perceiving Loudness, Pitch, and Location
160
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
160
Active Filters
673
Active filters are electronic circuits that use operational amplifiers (op-amps), resistors, and capacitors to filter out unwanted frequency components from a signal. A first-order low-pass active filter is designed to pass signals with a frequency lower than a certain cutoff frequency and attenuate frequencies higher than that cutoff frequency. The transfer function for a first-order low-pass active filter is:
673
Perception of Sound Waves
4.4K
The human ear is not equally sensitive to all frequencies in the audible range. It may perceive sound waves with the same pressure but different frequencies as having different loudness. Moreover, the perception of sound waves depends on the health of an individual's ears, which decays with age. The health of one's ears may also be affected by regular exposure to loud noises.
The pitch of a sound depends on the frequency and the pressure amplitude of the source. Two sounds of the same...
The pitch of a sound depends on the frequency and the pressure amplitude of the source. Two sounds of the same...
4.4K
Design Example
307
The innovation of touch-tone telephony revolutionized the telecommunications industry by replacing the traditional rotary dial with a dual-tone multi-frequency (DTMF) signaling system. This system uses a matrix-style keypad with buttons arranged in four rows and three columns, creating 12 distinct signals each assigned to a pair of frequencies. Each button press results in a simultaneous generation of two sinusoidal tones – one from a low-frequency group (697 to 941 Hz) and one from a...
307

