Related Experiment Video
Updated: Mar 2, 2026

Investigating the Three-dimensional Flow Separation Induced by a Model Vocal Fold Polyp
Published on: February 3, 2014
Interpretable vocal tract and respiratory inversion via physics-informed neural operators.
Mengtao Deng1, Cheng Liu2, Zhangmei Yang3
1Teacher Training College, Dazhou Vocational and Technical College, Dazhou, 635000, Sichuan, China. y05011817@163.com.
This study introduces a physics-informed neural network framework for accurate vocal tract modeling and real-time analysis. The new method enhances timbre fidelity and provides interpretable physiological insights for personalized voice applications.
Area of Science:
- Acoustics and Signal Processing
- Biomedical Engineering
- Machine Learning
Background:
- Physiological variations in vocal tracts and respiratory systems challenge accurate timbre modeling and real-time vocal analysis.
- Current data-driven methods often lack physical interpretability and speaker robustness.
- Accurate vocal tract geometry and respiratory dynamics reconstruction from acoustic data is crucial for advanced vocal analysis.
Purpose of the Study:
- To propose a physics-informed multimodal inversion framework using Kolmogorov-Arnold (KAN) operators for interpretable reconstruction of vocal tract geometry and respiratory dynamics.
- To achieve high timbre fidelity and low latency in vocal analysis.
- To provide a foundation for fine-grained timbre reconstruction and personalized vocal analysis.
Main Methods:
- A nested three-layer KAN inverts audio spectrum into vocal tract cross-sectional areas.
- A gated recursive module constrains pressure evolution using mass-momentum conservation.
- Super-resolution prediction head optimized with fractional-order temporal regularization and wave-equation residuals enhances timbre fidelity.
Main Results:
- Achieved log-spectral distortion of 1.83 ± 0.32 dB and sub-band error rate of 6.4 ± 1.1% (1.2-2.4 kHz).
- Demonstrated low end-to-end latency (14.2-18.3 ms) and compact memory usage (108-121 MB) on edge devices.
- Maintained minimal vocal tract geometry error (MAE-CSA ≤ 0.23) and respiratory estimation bias (RMSE-P ≤ 0.52) across unseen voice types.
Conclusions:
- Integrating neural operators with physical constraints enables accurate, interpretable, and real-time inversion of vocal physiology.
- The proposed framework offers a principled technical foundation for advanced timbre reconstruction.
- This approach paves the way for personalized vocal analysis and improved speech technology.
Related Concept Videos
Neural Control of Respiration
Respiratory Centers in the Brainstem
Two primary areas comprise the respiratory center: the medullary respiratory center in the medulla oblongata and the pontine respiratory group in the pons. The...
Anatomy of Respiratory System II: Lower Respiratory Tract
The Larynx
It is located between the pharynx and the trachea, acts as a passageway for air, and hosts several critical structures, such as the epiglottis, vocal cords, and glottis. The epiglottis acts as a gateway, guiding food to the...
Anatomy of Respiratory System I: Upper Respiratory Tract
Nose and nasal cavity
The nose and nasal cavity represent the main external openings of the respiratory tract....
Larynx
Anatomy of the Larynx
The larynx consists of various components, including cartilage, muscles, and vocal cords. Its structure includes three large unpaired cartilages—the thyroid, cricoid, and epiglottis—and three smaller paired cartilages—the arytenoids,...
Respiratory Volumes
Tidal Volume (TV) Tidal volume (TV) is the air inhaled or exhaled in a...
Application of Integration: Problem Solving

