Related Experiment Video
Updated: Aug 30, 2026

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
Published on: August 9, 2024
Visual Signal Typing in Voice Assessment: A Systematic Review of Methods, Reliability, and Clinical Applications
Prasanna Suresh Hegde1, Vijayalakshmi Subramaniam2, Vishal U S Rao1
1HealthCare Global Enterprises Ltd, Bengaluru, Karnataka, India.
Abstract:
Acoustic analysis is a fundamental step in the multidimensional assessment of dysphonia. However, conventional perturbation measures require nearly periodic signals for valid computation. Visual signal typing classifies voice signals into discrete types based on a narrowband spectrogram and serves as a critical gatekeeper for selecting appropriate analytical methods for a valid assessment. Two main frameworks exist: the Titze/Sprecher system for laryngeal voice and the van As-Brooks system for tracheoesophageal speech. Despite three decades of development, no systematic review had synthesized the conceptual and empirical evidence on the utility of visual signal typing. This systematic review followed PRISMA 2020 guidelines. PubMed, Scopus, Web of Science, and Google Scholar were searched from inception to early 2026. Studies were included if they used visual or acoustic signal typing as a primary or one of the methods with reported reliability, validity, classification accuracy, or clinical utility. Two independent reviewers performed screening, data extraction, and quality appraisal. Twenty-one studies (1995-2025) met inclusion criteria. The Titze/Sprecher laryngeal framework predominated in 15 studies, while the van As-Brooks framework was used in six tracheoesophageal (TE) studies. Inter-rater reliability ranged from κ = 0.45 to 0.95, with higher agreement in laryngeal studies (typically κ ≥ 0.72 with training) compared with TE studies (κ = 0.55-0.84). Two studies reported numeric perceptual coefficients (r = 0.746 and r = -0.93); the remaining studies reported moderate-to-strong directional associations. Automated approaches achieved 82-96% accuracy in classification of these signals. Interestingly, the signal typing process identified >50% of dysphonic samples as unsuitable for conventional perturbation analysis. Visual signal typing demonstrates moderate-certainty reliability and low-to-moderate-certainty validity as a clinical gatekeeper, particularly within the Titze/Sprecher laryngeal framework. Evidence is strongest for laryngeal voice with structured rater training. Significant gaps persist in clinical utility, deep learning automation, neurological voice disorders, connected speech, telemedicine, and international standardization.
