Related Experiment Videos
Machine Learning for Autism Spectrum Disorder Prediction: A Review of Data Augmentation and Feature Selection
Sahar Alkhaibari1,2, Feng Dong1
1Department of Computer and Information Sciences University of Strathclyde Glasgow United Kingdom.
Abstract:
Autism spectrum disorder (ASD) is a complex neurodevelopmental condition characterized by persistent difficulties in social communication, social interaction, and repetitive behaviors. Early and accurate diagnosis is essential but is often hindered by subjective clinical assessments, limited data availability, and inconsistencies in existing diagnostic tools. This review evaluates the role of machine learning and deep learning approaches in improving ASD prediction, with a particular focus on two important yet relatively underexplored methodological components: data augmentation and feature selection. A structured literature search was conducted across major scientific databases, including IEEE Xplore, PubMed, Scopus, and Google Scholar, to identify studies published between 2021 and 2024. The methodological quality and risk of bias of the included studies were assessed using the Prediction Model Risk of Bias Assessment Tool. A total of 26 peer-reviewed studies were included based on their relevance to machine learning/deep learning-based ASD prediction and their explicit application of data augmentation or feature selection techniques. Data augmentation methods were categorized into conventional approaches, such as geometric and color-space transformations, and advanced techniques, including generative adversarial network-based synthetic data generation. Although augmentation techniques may improve model robustness and help address dataset scarcity, relatively few studies conducted ablation analyses to isolate the contribution of individual augmentation strategies. Feature selection approaches were classified into filter, wrapper, and embedded methods. Commonly used techniques included information gain, chi-square tests, recursive feature elimination, and elastic net regularization. While these methods may improve predictive performance and model interpretability, they are frequently applied without sufficient empirical justification or biological interpretation. Overall, this review highlights important methodological limitations, including limited external validation, insufficient ablation analyses, and inadequate evaluation frameworks, which reduce confidence in reported performance improvements and model generalizability. Future research should emphasize methodological transparency, robust validation strategies, multimodal data integration, and clinically interpretable modeling approaches to improve the reliability and clinical applicability of ASD prediction systems.