Related Experiment Video
Updated: Feb 18, 2026

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
The Acoustic Voice Quality Index Version 03.01 for the Japanese-speaking Population
Kiyohito Hosokawa1, Ben Barsties V Latoszek2, Toshihiko Iwahashi3
1Department of Otorhinolaryngology, Japan Community Health care Organization (JCHO) Osaka Hospital, Osaka, Japan; Department of Otorhinolaryngology, Osaka Police Hospital, Osaka, Japan; Department of Otorhinolaryngology and Head & Neck Surgery, Osaka University Graduate School of Medicine, Osaka, Japan.
Objectives:
We aimed to determine the most appropriate syllable number for analyzing the Acoustic Voice Quality Index for the Japanese-speaking population (AVQIv3-JP) and to validate AVQIv3-JP using the determined syllable number.
Methods:
First, we counted how many syllables should be included in each continuous speech (CS) sample to achieve time-balanced analysis between CS and sustained vowel samples using our previous dataset including 336 CS samples with 58 syllables. From the descriptive statistics of the counted syllable numbers, the most appropriate syllable number was identified. Subsequently, we performed validation procedures of AVQIv3-JP using our latest dataset including 455 recordings.
Results:
Thirty Japanese syllables were judged to be the most appropriate syllable number. The concurrent validity of the AVQIv3-JP using 30 syllables was confirmed by Spearman's rho of 0.873. Subsequently, the receiver operating characteristic analysis demonstrated the excellent discriminative capability of AVQIv3-JP, showing the area under the curve of 0.915. The AVQIv3's original threshold of 2.43 in the Dutch language corresponded to sensitivity and specificity of 64.6% and 97.3%, respectively. In the present study, a threshold of 1.41 achieved the best accuracy with balanced sensitivity and specificity of 84.4% and 85.6%, respectively. Furthermore, the 95th percentile of the control participants exhibited a threshold of 2.06, showing sensitivity and specificity of 72.1% and 93.8%, respectively, as well as reasonable positive and negative likelihood ratios of 11.7 and 0.298, respectively.
Conclusion:
The AVQIv3 using 30 Japanese syllables is a reliable measurement tool for estimating the severity of voice quality and detecting abnormal voices.
More Related Videos
Related Concept Videos
Physical Assessment of the Respiratory Tract IV: Auscultation
Breath Sounds
Breath sounds are categorized into vesicular, bronchovesicular, and bronchial.
Sound Intensity Level
The human ear can perceive an extensive range of sound intensity, necessitating the use of the logarithmic scale to define a physical quantity—the intensity level. It is a ratio of two intensities and...
Pulse amplitude and quality
A weak or absent pulse may indicate reduced cardiac output or poor left ventricular contraction, which can be signs of cardiovascular dysfunction or...
Assessment of the Cardiovascular System IV: Auscultation
Normal Heart Sounds
S1 (First Heart Sound)-
S1 is made by the closure of the mitral and tricuspid valves (atrioventricular valves), marking the beginning of systole.
S2 (Second Heart Sound)-
S2 is made by the closure of the aortic and pulmonic valves (semilunar valves), marking the end of the systole.
Sound Intensity
Auditory Perception

