Related Experiment Video
Updated: Jul 16, 2026

Use of a Video Scoring Anchor for Rapid Serial Assessment of Social Communication in Toddlers
Published on: March 14, 2018
Label leakage unmasked: a trustworthy-AI audit of autism screening models using the CLEAR-RD framework
Boulbaba Ben Ammar1,2, Walid Karamti1,2
1Department of Computer Science, College of Computer, Qassim University, Buraydah, Saudi Arabia.
Abstract:
Trustworthy pediatric screening AI needs more than accuracy: label validity, calibration, subgroup robustness, and clinical-scope limits. We propose Clear-Rd (Clinical Leakage Evaluation and Audit Routine for Rule-Derived labels), a five-stage framework-Configure, Lift, Evaluate, Audit, Report-and apply it to two public autism-screening datasets-Q-CHAT-10 toddler (N = 1054) and AQ-10 child/adolescent/adult (N = 6075)-merged into a unified cohort (N = 7129); prior studies report accuracies above 95%. Three configurations-S 1 (full, with score), S 2 (items, no score), and S 3 (demographics only)-separate label leakage from demographic signal. A leakage audit confirms the Q-CHAT-10 label is exactly determined by score>3, the AQ-10 label approximately by score≥6. Sixteen models from three families (classical, deep-tabular, and prompt-structure simulators) are evaluated under 5-fold stratified cross-validation. On the leakage-free S 3, Gradient Boosting attains ROC-AUC 0.766 ± 0.005 and all families converge within 0.04 ROC-AUC, indicating the binding constraint is the feature set, not model capacity. Calibration, subgroup-parity, and cost-sensitive analyses complete the audit. The results expose a gap between reported and clinically interpretable performance; Clear-Rd is a reusable template for rule-derived medical labels. No such model should be interpreted as predicting clinical autism diagnosis without external validation against gold-standard outcomes.