Related Experiment Video
Updated: May 5, 2026

09:09
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
1.1K
Evaluating automated evaluation systems for spoken English proficiency: An exploratory comparative study with human
1City Culture and Communication College, Suzhou City University, Suzhou, Jiangsu Province, China.
Plos One
|March 28, 2025
Summary
Automated evaluation systems (AESs) for spoken language assessment show promise in English as a Foreign Language (EFL) contexts. Two out of three Chinese-developed AES tools aligned well with human ratings, suggesting potential for efficient language assessment.
Area of Science:
- Educational Technology
- Language Assessment
- Artificial Intelligence in Education
Background:
- Automated evaluation systems (AESs) are gaining traction globally for spoken language assessment.
- The validity of AESs in non-Western educational contexts, specifically for English as a Foreign Language (EFL) learners, is not well-established.
- Chinese-developed AES tools are widely used but require validation for assessing spoken English proficiency.
Purpose of the Study:
- To investigate the validity of three Chinese-developed AES tools in assessing spoken English proficiency.
- To compare AES scores with human ratings among Chinese undergraduate students.
- To address the underexplored area of AES validity in non-Western educational settings.
Main Methods:
- An IELTS-adapted speaking test was administered to 30 Chinese undergraduates.
- Spoken English proficiency was assessed simultaneously by three Chinese-developed AESs and human raters.
- Scoring alignment was analyzed using intra-class correlation coefficients, Pearson correlations, and linear regression.
Main Results:
- Two of the three AES tools demonstrated strong agreement with human rater scores.
- One AES exhibited systematic score inflation, indicating potential algorithmic limitations.
- Algorithmic discrepancies and insufficient consideration of linguistic nuances may explain the observed score inflation.
Conclusions:
- AESs can serve as effective complements to traditional language assessment methods in EFL contexts.
- Calibration and rigorous validation are crucial for ensuring the reliability and fairness of AESs.
- AESs offer potential for enhancing efficiency and standardization in spoken language assessment.

