Related Experiment Video
Updated: Oct 8, 2026

Advancing Dyslexia Assessment in Children Through Computerized Testing
Published on: August 16, 2024
Exploring the structural component of construct validity of integrated speaking tasks: The case of a large-scale
Yan Zhou1, Ke Bin2, Lawrence Jun Zhang3
1School of Foreign Studies, Southern Medical University, China; Center for International Communication of Healthy China and TCM Culture, Social Science Research Base of Guangdong Province, China; Faculty of Arts and Education, University of Auckland, New Zealand.
Abstract:
Integrated speaking tasks have been widely used in many large-scale high-stakes tests. However, little is known about their application among second language learners of English in K-12 learning contexts, such as in the Guangdong Version of the Computer-based English Listening and Speaking Test (hereafter, CELST) of the National Matriculation English Test. This misalignment is particularly problematic given the substantial impact of integrated speaking tasks on English language teaching, learning, and assessing in China and internationally. Employing a bi-factor Exploratory Structural Equation Modeling (ESEM) and hierarchical multiple regression analyses, the present study examined the structural component of CELST's construct validity by investigating whether the observed internal structure of test performance data provides empirical evidence for the hypothesized construct underlying the test, as manifested in its official test specifications issued by the Education Examinations Authority of Guangdong Province. Findings provided partial evidence that the internal structure of CELST performance data is consistent with the hypothesized construct, which involves students' ability to accomplish specific tasks in particular contexts by drawing on and applying multiple knowledge resources, including task prompts, knowledge of English and the world, source materials, and communicative strategies. Specifically, parameters representing the five hypothesized domain-specific language competence factors (with coherence represented just by one single indicator) and two hypothesized communicative strategies, extracted from the participants' oral output, jointly contributed to the participants' test performance, with weights differed across factors. This variability highlights the comprehensiveness and contextual specificity of the test. These findings could provide empirical evidence supporting the validity of score interpretations and offer important implications for the teaching, learning, and assessment of integrated speaking tasks in senior high schools in China.