Related Experiment Video
Updated: Sep 28, 2026

Multimodal Protocol for Assessing Metacognition and Self-Regulation in Adults with Learning Difficulties
Published on: September 27, 2020
Conditional validity in LLM-mediated L2 assessment: an argument-based systematic review and meta-analysis
Latifah Hamdan Alghamdi1, Talal Musaed Alghizzi2
1ELC, King Khalid University, Abha, Saudi Arabia.
Introduction:
Large language models (LLMs) are increasingly used for scoring and feedback in second-language (L2) assessment, yet the validity of the resulting interpretations remains contested. This review evaluated when LLM-mediated assessment is psychometrically and educationally defensible using an argument-based validity framework focused on construct representation, reliability and reproducibility, fairness, and washback.
Methods:
Following PRISMA 2020, we systematically searched nine databases and repositories for empirical studies published from January 2022 to December 2025. Fifty-two studies met the inclusion criteria. Quantitative synthesis used REML random-effects models where effects were sufficiently comparable; Pearson correlations, rank correlations, ICC/κ/QWK, and other psychometric indices were otherwise retained on metric-appropriate scales. Washback outcomes were synthesized using Hedges' g.
Results:
Source-level auditing showed that directly comparable Pearson evidence was sparse. Three source-verified L2 studies yielded moderate human-LLM score correspondence (r = 0.66, 95% CI [0.53, 0.75]) with substantial heterogeneity (I 2 = 70.8%). Evidence was strongest in rubric-guided, structurally constrained, and human-supervised contexts, whereas construct representation, fairness, operational reproducibility, and generalizability remained underdeveloped. Across 14 studies, LLM-mediated formative feedback produced a small-to-moderate positive effect on revision quality and short-term writing outcomes (g = 0.42, 95% CI [0.27, 0.57]), although heterogeneity was substantial (I 2 = 69.8%) and concerns about cognitive offloading, learner agency, and reproducibility persisted.
Discussion:
The findings support a context-dependent interpretation of validity rather than a universal claim about LLM assessment capability. Current evidence supports carefully bounded formative and human-supervised applications but does not justify autonomous high-stakes scoring or broad claims of construct validity, fairness, and generalizability.
