Related Experiment Videos
Revisiting judging reliability in taekwondo freestyle Poomsae: implications for AI-supported evaluation
Min-Woo Jeon1, Hong-Suk Kim2, Ji-Yong Park2
1Department of Taekwondo, College of Physical Education, Kyung Hee University, Yongin-si, Republic of Korea.
Background:
Before artificial intelligence-based judging systems can be applied to Taekwondo Poomsae, it is necessary to understand which performance components human judges can identify consistently and which remain difficult to evaluate reliably. This study examined reliability in freestyle Poomsae judging to identify evaluation components with stable and unstable human scoring.
Methods:
Ten internationally certified Taekwondo Poomsae referees evaluated ten official competition videos twice, separated by a one-week washout period. Systematic session effects were evaluated with separate linear mixed-effects models estimated by restricted maximum likelihood using the SPSS MIXED procedure. Session was specified as a fixed effect, judge and video as crossed random intercepts, and the two observations within each judge-video pair as repeated measurements with an unstructured residual covariance matrix. Fixed-effect inference used Satterthwaite denominator degrees of freedom. Test-retest stability was evaluated for identical judge-video pairs, with primary emphasis on the absolute-agreement ICC because systematic session effects were detected. Inter-rater agreement was evaluated separately by session using two-way random effects, single rating absolute-agreement and consistency ICCs with 95% confidence intervals and Kendall's W. Item-specific analyses were exploratory, and no multiplicity adjustment was applied.
Results:
The total score increased by 0.304 points, 95% CI [0.187, 0.421], t (99) = 5.177, p < 0.001. At the nominal 0.05 level, exploratory item-specific increases were observed for jumping side kick (β = 0.057, p < 0.001), acrobatic kicking technique (β = 0.021, p = 0.003), and expression of energy (β = 0.047, p = 0.003); the increase for basic movements and practicability did not reach the 0.05 threshold (p = 0.053). Absolute-agreement test-retest ICCs ranged from 0.039 to 0.848 and were excellent for harmony and the total score. All single-rating inter-rater ICC(A,1) point estimates were below 0.40. For the total score, ICC(A,1) was -0.007, 95% CI [-0.018, 0.036], in the first session and 0.060, 95% CI [0.010, 0.224], in the second session.
Conclusion:
Systematic session effects, absolute test-retest agreement, relative consistency over time, and inter-rater agreement represent distinct aspects of judging reliability. Evaluation components with poor single-judge agreement require clearer operational definitions and further validation before computational or AI-assisted scoring is considered.