Assessing different cross-validation schemes for predicting novel traits using sensor data: An application to dry
A Yilmaz Adkinson1, M Abouhawwash2, M J VandeHaar2
1Department of Animal Science, Michigan State University, East Lansing, MI 48824; Department of Animal Science, Erciyes University, 38039 Kayseri, Türkiye.
Journal of Dairy Science
|June 14, 2024
Summary
Milk mid-infrared (MIR) spectral data can predict feed efficiency in dairy cows when validated within cow groups. However, predictive accuracy significantly decreases in herd-independent validations, highlighting the need for improved algorithms and calibration.
Area of Science:
- Animal Science
- Dairy Production
- Genetics and Genomics
Background:
- Feed efficiency is crucial for dairy farm profitability, but daily dry matter intake (DMI) recording is costly.
- Mid-infrared (MIR) spectral data from milk offers a potential proxy for DMI prediction.
- Accurate DMI prediction is essential for genetic selection and improving feed efficiency.
Purpose of the Study:
- To evaluate the utility of milk MIR spectral data for predicting proxy phenotypes of DMI.
- To compare predictive models including MIR data, energy sinks, or both, across different cross-validation schemes.
- To assess the performance of MIR-based DMI prediction under cow-independent, experiment-independent, and herd-independent scenarios.
Main Methods:
- Developed three models: MIR data only (M1), energy sinks (body weight, weight change, milk energy) plus cow-level variables (M2), and M2 plus MIR data (M3).
- Utilized weekly and 28-day DMI records from US Holstein cows across multiple experiments and herds.
- Employed 10-fold cow-independent, 10-fold experiment-independent, and 4-fold herd-independent cross-validation schemes.
Main Results:
- Cow-independent cross-validation showed significant improvement with MIR addition (M3 vs. M2), reducing RMSE to 1.59 kg and increasing R² to 0.89.
- Experiment-independent and herd-independent cross-validations revealed no significant benefit of MIR inclusion (M3 vs. M2).
- Broader cross-validation schemes (experiment- and herd-independent) showed increased bias, with herd-independent validation being most critical for real-world application.
Conclusions:
- Milk MIR data shows potential for predicting DMI proxy phenotypes, but its effectiveness is highly dependent on the cross-validation strategy.
- Current MIR-based prediction models lack robustness for herd-independent applications, which are essential for widespread genetic evaluations.
- Future research should focus on developing algorithms suitable for broader cross-validation and improving spectrophotometer calibration for reliable DMI proxy prediction.


