Related Experiment Video
Updated: Jun 7, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Can acoustic measurements predict gender perception in the voice?
Diego Henrique da Cruz Martinho1, Leonardo Wanderley Lopes2, Rodrigo Dornelas3
1Department of Human Development, Universidade Estadual de Campinas-UNICAMP, Campinas, Brazil.
Purpose:
To determine if there is an association between vocal gender presentation and the gender and context of the listener.
Method:
Quantitative and transversal study. 47 speakers of Brazilian Portuguese of different genders were recorded. Recordings included sustained vowel emission, connected speech, and the expressive recital of a poem. Subsequently, four scripts were used in Praat to extract 16 acoustic measurements related to prosody. Voices underwent Auditory-Perceptual Assessment (APA) of the gender presentation by 236 people [65 speech and language pathologist (SLP) with experience in the area of the voice (SLP), 101 cisgender people (CG), and 70 transgender and non-binary people (TNB)]. Gender presentation was evaluated by visual analogue scale. Agreement analyses were executed among quantitative variables and multiple linear regression models were generated to predict APA, taking the judge context/gender and speaker gender into consideration.
Results:
Acoustic analysis revealed that cis and transgender women had higher median fundamental frequency (fo) values than other genders. Cisgender women exhibited greater breathiness, while cisgender men showed more vocal quality deviations. In terms of APA, significant differences were observed among judge groups: SLP judged vowel samples differently from other groups, and TNB judged speech samples differently (both p<0.001). The predictive measures for the APA varied based on the sample type, speaker gender, and judge group. For vowel samples, only SLP judges had predictive measures (fo and ABI Jitter) for cisgender speakers. In number counting samples, predictive measures for cisgender speakers included fomed and HNR for CG judges, and fomed for both SLP and TNB judges. For transgender and non-binary speakers, predictive measures were fomed for CG and SLP judges, and fomed, CPPs, and ABI for TNB judges. In the poem recital task, predictive measures for cisgender speakers were fomed and HNR for both SLP and CG judges, with additional measures of cvint and sr for CG judges, and fomed, HNR, cvint, and fopeakwidth for TNB judges. For transgender and non-binary speakers, the predictive measures included a wider range of acoustic features such as fomed, fosd, sr, fomin, emph, HNR, Shimmer, and fo peakwidth for SLP judges, and fomed, fosd, sr, fomax, emph, HNR, and Shimmer for CG judges, while TNB judges considered fomed, sr, emph, fosd, Shimmer, HNR, Jitter, and fomax.
Conclusions:
There is an association between the perception of gender presentation in the voice and the gender or context of the listener and the speaker. Transgender and non-binary judges diverged to a higher degree from cisgender and SLP judges. Compared to the evaluation of cisgender speakers, all judge groups used a greater number of acoustic measurements when analyzing the speech of transgender and non-binary individuals in the poem recital samples.
Related Concept Videos
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Factors Affecting Perception
An illustrative example of a perceptual set is the scenario where an airline pilot told...
Socioemotional Experience and Gender Development
Perception of Sound Waves
The pitch of a sound depends on the frequency and the pressure amplitude of the source. Two sounds of the same...

