Restoration of Acoustic Identity via Artificial Intelligence-Driven Neural Voice Conversion for Total Laryngectomy
Mehmet Can Girgin1, Cagdas Can2, Pınar Onucak Girgin3
1Department of Emergency Medicine, Istanbul Beykent University, Istanbul, TUR.
Abstract:
Total laryngectomy results in the permanent loss of natural phonation, necessitating a shift from functional speech restoration to the holistic recovery of acoustic identity. This technical report presents a framework for artificial intelligence (AI)-driven Neural Voice Conversion designed to transform mechanical or esophageal speech into a patient's unique preoperative voice. The technical architecture is centered on a Non-Parallel Many-to-One Voice Conversion system, specifically utilizing a Generative Adversarial Network backbone, such as CycleGAN-VC3, combined with a Variational Autoencoder for latent feature disentanglement. The framework begins with the extraction of fundamental frequency ($F_0$) and spectral envelopes from the patient's postoperative speech using high-resolution vocoders such as World or WaveNet. To ensure biometric fidelity, a neural refinement layer processes Mel-frequency cepstral coefficients to match the unique timbre and glottal flow characteristics stored in the patient's preoperative digital voice bank. A critical innovation in this framework is the inclusion of a "Biometric Integrity Module," which ensures the synthesized output maintains a high enough Mel-spectrogram resolution to satisfy the equal error rate requirements of voice-activated security systems. Furthermore, the system employs a Speaker-Encoder network that maps voice identity into a high-dimensional embedding space, allowing the model to preserve prosodic nuances while eliminating the robotic artifacts typical of traditional electrolarynx devices. By integrating a Pitch-Contour Alignment algorithm, the framework synchronizes the emotional intent of the speaker with the regenerated acoustic signal. This technical approach addresses not only the phonetic intelligibility of speech but also the preservation of the patient's digital persona and social inclusion. The report outlines the multi-stage training process involving phonetic loss functions and adversarial training to minimize distortion. By bridging the gap between surgical outcome and biometric security, this AI-driven framework redefines postoperative rehabilitation as the restoration of the patient's complete acoustic and digital identity.

