Related Experiment Videos
Explainable machine learning for age-specific talent identification in Chinese youth football
Background And Aims:
To identify elite youth football players in China, we developed an interpretable machine learning model.
Methods:
We modeled each age group separately to track changes in predictive factors across different age groups. A cross-sectional analysis included 398 players from the 2024 Yunnan Football Association Elite Training Camp (U8 to U12). All participants underwent a multidimensional assessment covering anthropometric, physical, technical, psychological, and match performance dimensions. In our prediction model, coaches' selection of provincial elite players was used as the outcome variable. Three classifiers: XGBoost, Random Forest, and Logistic Regression, were trained with hyperparameter tuning and 5-fold cross-validation. Then we used SHapley Additive exPlanations (SHAP) to interpret our models.
Results:
The best-performing model demonstrated high discriminative ability (test set ROC-AUC: 0.900-0.957). SHAP analysis indicated that dribbling speed and standing long jump consistently mattered at all ages, while sprint performance were especially important for the U8 and U10 groups. Notably, for U12 males, left-foot shooting skill was the strongest predictor, while for U12 females, tactical awareness during matches emerged as most important. Psychological traits contributed little predictive value. When we applied retrospectively to national selection data, the U12M model showed fair agreement (K = 0.269) but identified 40%-47% of national elites and highlighted several overlooked high-potential players.