机器学习算法和多重线性回归的比较,用于估计阿卡拉曼羔羊的活体重
Özge Kozaklı1, Ayhan Ceyhan2, Mevlüt Noyan3
1Department of Animal Production and Technologies, Faculty of Agricultural Sciences and Technologies, Niğde Ömer Halisdemir University, Niğde, 51240, Turkey. ozgekozakli94@hotmail.com.
Tropical animal health and production
|September 3, 2024
概括
机器学习算法,特别是随机森林,准确地预测阿卡拉曼羊羔断奶后的体重. 这项研究比较了各种模型,发现随机森林优于传统方法,例如用于牲畜体重预测的多重线性回归.
科学领域:
- 农业科学 农业科学
- 动物科学动物科学
- 数据科学与机器学习
背景情况:
- 准确预测牲畜的断奶后体重对于有效的农场管理和经济可行性至关重要.
- 传统的统计方法可能无法完全捕捉影响动物生长的复杂关系.
- 阿卡拉曼品种是一个重要的土耳其羊品种,使其增长预测对区域农业很重要.
研究的目的:
- 用多重线性回归和各种机器学习算法预测阿卡拉曼羔羊的断奶后体重.
- 基于诸如水年龄,性别和出生体重等因素的不同算法的预测性能进行比较.
- 确定在现实世界农业环境中预测羔羊体重的最准确模型.
主要方法:
- 分析了来自多个农场的25,316只阿卡拉曼羊羔的数据.
- 测试的算法包括多重线性回归,随机森林,支持矢量机器,XGBoost和各种神经网络.
- 使用K折交叉验证来评估模型性能,使用调整R平方,RMSE,MAD和MAPE指标.
主要成果:
- 随机森林算法展示了最高的预测性能,达到0.75.75的调整R平方.
- 随机森林的RMSE,MAD和MAPE值分别为3.683,2.876和10.112,这表明它的准确性很强.
- 人工神经网络分析通常比多重线性回归提供了比多重线性回归更准确的活体重预测.
结论:
- 与传统方法相比,机器学习模型,特别是随机森林,在预测阿卡拉曼羊羔断奶后体重方面提供了更高的准确性.
- 这项研究证实了先进的计算技术在优化畜牧管理和育种策略方面的潜力.
- 性能标准的低标准偏差表明这些模型是稳固的,没有过度装配问题.
相关概念视频
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
45
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
45
Residuals and Least-Squares Property
7.3K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.3K
Microsoft Excel: Regression Analysis
533
Regression analysis in Microsoft Excel is a powerful statistical method for examining the relationship between a dependent variable and one or more independent variables. It's used extensively in fields such as economics, biology, and business to predict outcomes, understand relationships, and make data-driven decisions. The most common type is linear regression, which attempts to fit a straight line through the data points to model the relationship between variables.
To perform regression...
To perform regression...
533
Correlation and Regression
1.2K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
1.2K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K


