Related Experiment Video
Updated: Jun 22, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
LUMINA: Linguistic unified multimodal Indonesian natural audio-visual dataset
Eka Rahayu Setyaningsih1,2, Anik Nur Handayani1, Wahyu Sakti Gunawan Irianto1
1Department of Electrical Engineering and Informatics, Universitas Negeri Malang, Semarang Street 5, Malang, 65145, East Java, Indonesia.
Abstract:
The LUMINA (Linguistic Unified Multimodal Indonesian Natural Audio-Visual) Dataset is a carefully curated constrained audio-visual dataset designed to support research in the field of speech perception. Spoken exclusively in Indonesian, LUMINA contains high-quality audio-visual recordings featuring 14 native speakers, including 9 males and 5 females. Each speaker contributes approximately 1,000 sentences, producing a rich and diverse data collection. The recorded videos focus on facial recordings, capturing essential visual cues and expressions that accompany speech. This extensive dataset provides a valuable resource for understanding how humans perceive and process spoken language, paving the way for speech recognition and synthesis technology advancements.
More Related Videos
Related Concept Videos
Lateralization
Auditory Pathway
When viewed cross-sectionally, the cochlea reveals the scala vestibuli and scala tympani flanking...
Light Acquisition
Deconvolution
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
Intralumenal Vesicles and Multivesicular Bodies
Upsampling

