Related Experiment Video
Updated: Aug 6, 2026

Instrumentation of Near-term Fetal Sheep for Multivariate Chronic Non-anesthetized Recordings
Published on: October 25, 2015
Machine learning based identification of key production drivers of sheep population in Türkiye: a century-long
Malik Ergin1, Merve Mürüvvet Dağ2, Bektaş Kadakoğlu2
1Department of Animal Science, Faculty of Agriculture, Isparta University of Applied Sciences, Isparta, Türkiye. malikergin@isparta.edu.tr.
Abstract:
For the first time, this study employs a century-long dataset (1925-2024) to reveal key factors that would be in relationship with sheep population (NSheep) in Türkiye using state-of-the-art machine learning algorithms. Due to the existence of missing values in the original dataset, missing observations were addressed through four imputation techniques-Next Observation Carried Backward (NOCB), Mean, MIDASpy, and Random Forest (RF)-generating four distinct datasets for comparative analysis. For revealing the key production factors related with NSheep, Extreme Gradient Boosting (XGB) and Multilayer Perceptron (MLP) algorithms were modeled via 5-fold cross-validation and multiple performance metrics (R², MSE, RMSE, MAE, and MdAPE). MLP produced lower prediction errors than XGB across all imputation techniques, though this difference was statistically confirmed only under NOCB and RF imputation (Diebold-Mariano test, P < 0.01 and P < 0.05, respectively); differences under MEAN and MIDASpy imputation were not significant. The highest overall accuracy was achieved by MLP with NOCB imputation (R² = 0.975), while XGB with RF imputation showed the weakest fit (R² = 0.917). Feature importance analyses consistently identified cattle population (NBovine) as the dominant variable associated with NSheep across all four imputation techniques and both algorithms, followed by meadow and pasture area (M&PH) for XGBoost and a more evenly distributed set of variables (M&PH, sheep meat production, goat population) for MLP. Given that NBovine and NSheep both increased steadily over the study period, this association is interpreted as reflecting shared structural growth among livestock subsectors rather than a causal effect of cattle population on sheep numbers. These dataset-specific, associative findings support the value of combining multiple imputation strategies with flexible machine learning algorithms to characterize structural interdependencies in long-term agricultural production data. Future studies incorporating chronologically ordered validation schemes and explicitly modeling structural breaks and policy shifts could further clarify the robustness of these associations, and extending this approach to other livestock species and regions would help establish their broader generalizability.
Related Concept Videos
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Cloning of Dolly the Sheep