Related Experiment Videos
CiwGAN and fiwGAN: Encoding information in acoustic data to model lexical learning with Generative Adversarial
1Department of Linguistics, University of California, Berkeley, United States of America.
Summary
This study introduces novel deep neural network architectures, Categorical InfoWaveGAN and Featural InfoWaveGAN, for unsupervised lexical learning from raw acoustic data. These models demonstrate emergent phonetic and phonological representation, enabling novel word generation and offering insights into speech processing.
Area of Science:
- Artificial Intelligence
- Cognitive Science
- Speech Processing
Background:
- Deep neural networks (DNNs) struggle to encode human speech information into raw acoustic data.
- Unsupervised learning of lexical information from audio is a significant challenge in speech processing.
- Existing Generative Adversarial Network (GAN) architectures require adaptation for complex audio data modeling.
Purpose of the Study:
- To propose novel DNN architectures for unsupervised lexical learning from raw acoustic inputs.
- To model how DNNs can encode linguistic information into acoustic representations.
- To investigate the emergent properties of these models, including novel word generation and latent space interpretability.
Main Methods:
- Developed two GAN-based architectures: Categorical InfoWaveGAN (ciwGAN) and Featural InfoWaveGAN (fiwGAN).
- Combined Deep Convolutional GAN (WaveGAN) with InfoGAN for audio data processing.
- Introduced a new latent space structure for simultaneous featural and categorical learning, enabling low-dimension vector representations of lexical items.
- Trained networks on the TIMIT corpus for lexical item encoding and retrieval of latent codes from generated audio.
Main Results:
- Networks successfully encoded unique information corresponding to lexical items in their latent space as categorical variables.
- Manipulation of latent variables allowed the generation of specific lexical items.
- Networks occasionally generated innovative, linguistically interpretable lexical items not present in training data, demonstrating productive recombination of phonetic and phonological representations.
- Probing latent featural codes beyond training range resulted in prototypical lexical item generation, revealing underlying code values.
Conclusions:
- The proposed architectures enable unsupervised lexical learning from raw acoustic data, mirroring human speech productivity.
- The models offer a framework for understanding DNNs' latent space interpretability and meaningful representation learning.
- This research has potential applications in unsupervised text-to-speech generation within the GAN framework.