Related Experiment Videos
Fairness-aware artificial intelligence tutoring for multilingual learners: evaluating adaptive feedback across L1 and
Patrick O Akinwumi1, Meihua Qian1
1College of Education, Clemson University, Clemson, SC, United States.
Introduction:
Recent advances in large language models (LLMs), such as GPT-4, have created new opportunities for intelligent computer-assisted language learning, particularly through personalized feedback for multilingual English learners. Building on research in automated writing evaluation (AWE), this study developed an equity-aware AI tutoring framework that integrates GPT-generated feedback with rule-based simulated feedback for learner spelling errors. The framework draws on annotated TOEFL-Spell data enriched with learner first-language (L1) and English proficiency information to examine feedback quality, accessibility, personalization, and potential subgroup disparities.
Methods:
A verified matched analysis of 32 complete GPT-simulated feedback pairs was conducted using feedback length, Flesch Reading Ease, BERTScore, non-parametric subgroup tests, paired comparisons, effect-size estimates, and exploratory regression models controlling for selected input characteristics.
Results:
GPT-generated feedback demonstrated substantial semantic correspondence with simulated feedback, achieving an overall BERTScore precision of 0.8370, recall of 0.8742, and F1 score of 0.8552. GPT feedback was significantly longer than simulated feedback (M = 608.12 vs. 218.69 characters, p < 0.001) and, contrary to the preliminary analysis, was also significantly more readable on average (M = 75.62 vs. 64.64, p < 0.001). GPT feedback length did not differ significantly across L1 (p = 0.143) or proficiency groups (p = 0.422), while GPT readability showed a significant omnibus difference across proficiency levels (p = 0.047); however, no pairwise comparison remained significant after Holm correction. BERTScore F1 was stable across both L1 (p = 0.775) and proficiency groups (p = 0.500). Exploratory controlled analyses likewise provided no consistent evidence that L1 independently predicted feedback readability or semantic alignment after accounting for prompt and error characteristics.
Discussion:
Accordingly, observed subgroup variation is interpreted as descriptive heterogeneity rather than definitive evidence of algorithmic or sociolinguistic bias. Because L1 and proficiency were not fully crossed in the analytical sample, their independent effects could not be completely disentangled. A Gradio-based interface (https://huggingface.co/spaces/Oluyori/equity-ai-tutor) was additionally developed to demonstrate the practical deployment of the tutoring framework. The findings demonstrate LLM-based feedback potential for semantically aligned and accessible language support while emphasizing the need for larger, balanced samples, learner-centered evaluation, and multidimensional fairness assessment before claims regarding equitable educational performance can be made.