Related Experiment Video
Updated: Sep 2, 2026

Use of a Video Scoring Anchor for Rapid Serial Assessment of Social Communication in Toddlers
Published on: March 14, 2018
Reliability and Concurrent Validity of the Likert-Type Stuttering Severity Scale and the Visual Analog Scale in
Ayşe İlayda Mutlu1, İlkem Uçal2
1Speech and Language Therapy Department, Faculty of Health Sciences, Başkent University, Ankara, Turkiye.
Purpose:
This study examined the test-retest and inter-rater reliability of non-clinician severity judgments using two brief rating tools-the 11-point Severity Scale (SEV) and the Visual Analog Scale (VAS)-and evaluated their concurrent validity with clinician-rated Stuttering Severity Instrument-4-Turkish version (SSI-4-TR) scores.
Method:
This observational cross-sectional study used speech samples from 15 adults who stutter as stimulus materials. Recordings were evaluated by two speech-language therapists (SLTs) and two independent groups of non-clinician raters (n = 26 per group). SLTs assessed stuttering severity using the SSI-4-TR. One non-clinician group rated severity using a VAS, whereas the other used SEV. All ratings were completed online across three sessions. To examine temporal stability, 30% of the samples were re-rated after a one-week interval.
Results:
Both tools demonstrated excellent test-retest reliability, with intraclass correlation coefficients of 0.997 for the VAS and 0.987 for the SEV. Inter-rater reliability was moderate, with ICC values of 0.656 for the VAS and 0.573 for the SEV. Both scales showed strong positive correlations with SSI-4-TR scores, with correlation coefficients of 0.835 for the VAS and .814 for the SEV (p < 0.001).
Conclusions:
The findings indicate that the VAS and SEV demonstrated high test-retest reliability and moderate inter-rater reliability and showed strong associations with SSI-4-TR scores, supporting the use of brief perceptual rating scales as practical screening instruments in large-scale and time-limited clinical and research settings where rapid and resource-efficient assessment is required.
What This Paper Adds:
What is already known on this subject Perceptual severity ratings are widely used to estimate stuttering severity in both clinical and research contexts. Likert-type scales have demonstrated acceptable reliability when applied by trained listeners, whereas the Visual Analog Scale (VAS) has shown strong psychometric performance in several perceptual domains such as voice and resonance assessment. However, empirical evidence regarding the reliability and validity of a single-item VAS for evaluating stuttering severity remains limited. In addition, it is not yet clear whether trained non-clinician raters can apply brief perceptual severity scales consistently, particularly when ratings are compared with standardized clinical measures such as the SSI-4. What this study adds to the existing knowledge This study provides novel psychometric evidence comparing an 11-point Likert-type severity scale (SEV) and a Visual Analog Scale (VAS) for rating stuttering severity using trained non-clinician raters. Both tools demonstrated excellent test-retest reliability, moderate inter-rater reliability, and strong concurrent validity with clinician-rated SSI-4-TR scores. The findings extend previous research by demonstrating that brief perceptual severity ratings can produce stable and clinically meaningful estimates when administered by trained but non-expert listeners. The study also contributes to the literature by directly comparing continuous and categorical response formats within the same experimental framework. What are the clinical implications of this study? The findings suggest that brief perceptual severity scales such as VAS and SEV may serve as practical adjunct tools for screening, monitoring change, and large-scale data collection in clinical and research settings where time and resources are limited. Trained non-clinician raters may provide consistent evaluations under controlled conditions, supporting their potential role in structured assessment contexts. However, perceptual ratings should not replace comprehensive clinical evaluation and may be most informative when used alongside standardized measures such as the SSI-4 or frequency-based indices to obtain a more complete representation of stuttering severity.
