Related Experiment Video
Updated: Sep 5, 2026

An Efficient Single-Person Technique for Milk Sampling from Laboratory Mice
Published on: March 28, 2025
Explainable ensemble machine learning for pregnancy screening in dairy cows using routine dairy herd improvement
Byungho Chae1, In-Hyeok Cheon2, Nag-Jin Choi2
1Department of AI in Agriculture, Jeonbuk National University, Jeonju, 54896, Republic of Korea.
Abstract:
Early identification of non-pregnant dairy cows is important for minimizing days open, yet current pregnancy diagnosis methods require additional cost, equipment, or veterinary intervention. In this retrospective observational study, we developed an ensemble machine learning model to screen pregnancy status using routine Dairy Herd Improvement (DHI) test-day data together with the date of first insemination recorded in the same DHI database. Test-day records (n = 60,301) from 2,614 Holstein cows across 36 farms in South Korea were used to construct 50,826 observation windows, each comprising 3 consecutive monthly records. Seventy features were engineered: test-day milk traits, days in milk (DIM), and 2 reproductive-context features derived from the first-insemination date. Three base learners (random forest, XGBoost, logistic regression) were combined via soft voting. Under cow-level grouped 5-fold cross-validation (positive class = pregnant), the ensemble achieved an AUC of 0.857 ± 0.006, PR-AUC of 0.949, and a Brier score of 0.119. Leave-one-farm-out cross-validation confirmed generalization to unseen farms within the studied DHI system (AUC = 0.863 ± 0.034). In SHAP analysis, the 6 test-day milk trait groups together formed the largest domain (53.9% combined), reproductive context was the largest single feature group (26.3%), and DIM contributed 17.8% as a time proxy for lactation stage. Even when DIM, reproductive-context, and parity features were removed, a test-day milk traits model retained an AUC of 0.794, and DIM-stratified analysis showed lower milk yield in pregnant cows within every DIM interval (Cohen's d = -0.24 to -0.43). A dual-threshold strategy provided 2 operating points: an F1-maximizing threshold (sensitivity = 0.966, specificity = 0.391) for surveillance and a cost-minimizing threshold (sensitivity = 0.680, specificity = 0.844, precision = 0.934) for non-pregnancy screening. Because 3 mo of data accumulation are required, the model serves as a monthly safety net within existing DHI workflows for cows missed by conventional early pregnancy diagnosis.