Related Experiment Videos
Leakage-aware pre-event machine learning evaluation for horse-race prediction under temporal validation
1Graduate School of Psychology, Kansai University, Suita, Japan.
Abstract:
Horse-racing prediction was used as a case study to develop and evaluate a leakage-aware machine-learning framework based on information available after race-entry finalization and before outcomes were known. Japanese Racing Association flat-race data obtained through JRA-VAN Data Lab. and exported using TARGET frontier JV were split temporally: 2015-2022 for training, 2023-2024 for validation, and January 5, 2025 to May 10, 2026 for independent testing. The test set contained 63,910 horse-level observations from 4,556 races. Confidence intervals were estimated using 10,000 race-level bootstrap resamples. Two horse-level binary outcome labels were evaluated: a win outcome and a JRA place-rule-compatible place outcome derived from the processed field_size variable. Compared with the augmented current_full_model, the hyperparameter-matched no-theory sensitivity-analysis model matched_no_theory_model performed better for both outcomes. For the win outcome, ROC AUC was 0.7543 (95% CI: 0.7475-0.7609) for matched_no_theory_model and 0.7293 (95% CI: 0.7224-0.7362) for current_full_model. For the JRA place-rule-compatible place outcome, ROC AUC was 0.7513 (95% CI: 0.7469-0.7558) and 0.7164 (95% CI: 0.7118-0.7212), respectively. An exploratory race-level uncertainty diagnostic was evaluated separately and was not incorporated into the horse-level prediction models. These findings indicate that the evaluated theory-score block did not improve temporal test-set performance. Domain-derived feature blocks should be evaluated incrementally through leakage audits, temporal validation, feature-block ablation, and matched sensitivity analyses before deployment.