Related Experiment Video
Updated: Sep 23, 2026

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Predicting competitive outcomes in professional e-sports from 60-second pre-match voice acoustics using machine
Gabriel Kadri1, Raphael I M Santos1, Felipe O Aguiar1,2
1Mental Health Department, Santa Casa de São Paulo School of Medical Sciences, São Paulo, SP, Brazil.
Introduction:
Competitive performance in electronic sports (e-sports) depends on rapid decision-making, emotional regulation, and coordinated communication, yet little is known about whether pre-match vocal behavior contains information predictive of competitive outcomes.
Methods:
This study applied supervised machine learning to 68 acoustic features extracted at the frame level (50-ms frames, 25-ms step) from 60-second pre-match team communication recordings in 89 professional Counter-Strike: Global Offensive matches; frame-level predictions were aggregated into a match-level score representing the proportion of frames classified as a win. Three predictive conditions were evaluated, acoustic features only, ranking difference only, and a combined model integrating both, using stratified group five-fold cross-validation. Uncertainty was quantified using percentile 95% confidence intervals from 2,000 match-level bootstrap resamples of the pooled out-of-fold predictions, with chance-level discrimination defined as AUC = 0.50.
Results:
Across algorithms, models combining voice and ranking achieved the strongest performance, with the Decision Tree classifier reaching a mean bootstrap AUC of 77.3% (95% CI 65.8-86.9) and accuracy of 78.6% (95% CI 69.9-86.7); all five voice-plus-ranking models had 95% CIs excluding chance. Voice-only models showed more limited evidence of above-chance discrimination: only the Decision Tree (AUC 67.8%, 95% CI 54.8-79.2) and Random Forest (AUC 64.0%, 95% CI 50.7-76.3) had confidence intervals excluding 0.50, whereas Linear Discriminant Analysis, Logistic Regression, and k-Nearest Neighbors did not. No ranking-only model showed a confidence interval excluding chance. Exploratory LIME-based feature-attribution analyses indicated that ranking difference received the highest within-model attribution in the combined models, while delta spectral flux, chroma standard deviation, and spectral centroid received the highest within-model attribution among acoustic descriptors for tree-based, linear, and distance-based classifiers, respectively; these rankings are descriptive and were not subjected to formal cross-model statistical comparison.
Discussion:
These findings provide preliminary, dataset-bounded evidence that acoustic patterns in brief pre-match team communication were associated with match outcome and, for some models, contributed predictive information beyond ranking; the retrospective, single-team design does not establish a generalizable behavioral biomarker or a causal link between vocal acoustics and competitive readiness.

