Related Experiment Video
Updated: Sep 10, 2025

Synthetic, Multi-Layer, Self-Oscillating Vocal Fold Model Fabrication
Published on: December 2, 2011
A scalable codec for bone-conducted speech based on generative and diffusion models
Xiaolong Hu1, Zhe Chen1, Fuliang Yin1
1Department of Information and Communication Engineering, Dalian University of Technology (DUT), Dalian, 116023, China.
Abstract:
In extremely noisy communication scenarios, the bone-conducted microphone (BCM) speech codec is often combined with speech bandwidth extension to improve the BCM speech quality. However, this tandem approach leads to a complex system architecture. To address the problem, a scalable codec for BCM speech based on generative and diffusion probabilistic models is proposed in this paper. Specifically, a specialized codec architecture is constructed to encode BCM speech while complementing its high-frequency components. Then, a key feature extraction block is presented to address the diminishing memory capacity in shallow layers as the network depth increases. Next, considering the potential lack of high-frequency detail information, an overall refinement block is introduced to refine the reconstructed speech signals. Finally, based on the U-Net architecture, a diffusion probability model is proposed to upsample the input audio signal from a bandwidth of 8 kHz to a high-resolution audio signal with a bandwidth of 20 kHz and a sampling rate of 48 kHz. The proposed method can simultaneously encode and improve BCM speech quality using a single network. It supports different bitrate settings without architectural changes or retraining and dynamically adjusts transmitted data based on changing network load. Simulation experiments demonstrate its feasibility.
More Related Videos
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
06:24Author Spotlight: Advancements in the Fabrication of Synthetic Vocal Fold Models for Phonetic and Robotic Applications
Published on: January 5, 2024
Related Concept Videos
Bone Remodeling
Linear Approximation in Frequency Domain
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
Sampling Methods: Overview
In analytical chemistry, the choice of...
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Downsampling
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...