对缺少的乳牛生产数据的归算方法的比较
1Department of Animal Biosciences, University of Guelph, Guelph, ON, Canada.
Animal : an international journal of animal bioscience
|September 2, 2023
概括
随机森林归算有效地处理乳牛记录中缺少的数据,优于其他方法对缩DM摄入量,牛奶产量和体重. 这提高了现实世界动物数据集的价值.
科学领域:
- 动物科学 动物科学
- 数据科学数据科学数据科学
- 农业技术 农业技术
背景情况:
- 农场动物数据收集正在增加,产生了关于料摄入量,生长和环境影响的大量数据集.
- 现实世界的数据经常包含由于错误或错位而缺失的值,从而减少了数据的实用性.
- 归算对于解决缺失数据至关重要,但方法选择会影响分析结果.
研究的目的:
- 为了比较各种归算方法的性能,以估计乳牛数据集中的缺失值.
- 确定对诸如缩DM摄入量,料DM摄入量,牛奶产量和体重等变量最有效的归算技术.
主要方法:
- 清理了一大批奶牛数据集 (454,553条记录),结果有437,075条观察结果.
- 缺少的值存在于缩DM摄入量 (2.30%),料DM摄入量 (8.05%),牛奶产量 (15.12%) 和体重 (64.33%) 中.
- 在五个数据子集上评估了四种单变量和九种多变量归算方法,包括机器学习算法和MissForest.
主要成果:
- 随机森林归算在所有归算变量中显示出卓越的性能.
- 与其他方法相比,它实现了较低的平均平方预测误差和更高的一致性相关系数.
- 随机森林特别有效地证明了缩DM摄入量,牛奶产量和体重.
结论:
- 随机森林是解决大规模乳牛数据集中缺失数据的最佳归算方法.
- 使用随机森林的有效归算增强了真实世界动物生产数据的分析价值.
- 这一发现支持了数据质量的提高和精准畜牧业更可靠的洞察力.
相关概念视频
Mechanistic Models: Compartment Models in Individual and Population Analysis
64
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
64
Production Efficiency
16.9K
Net production efficiency (NPE) is the efficiency at which organisms assimilate energy into biomass for the next trophic level. Due to low metabolic rates and less energy spent on thermoregulatory processes, the NPE of ectotherms (cold-blooded animals) is 10 times higher than endotherms (warm-blooded animals).
16.9K
Cloning of Dolly the Sheep
3.9K
The first successfully cloned mammal was Dolly, a sheep, born on 5th July 1996 at Roslin Institute, Scotland. The cloned sheep was named after the American singer Dolly Parton. Dolly lived for seven years and died of respiratory complications, which is speculated to be due to the actual age of her DNA. Because the DNA in cloned cells belongs to an older individual, the cloned individual’s life expectancy may be affected. Indeed, analysis of Dolly’s DNA revealed shorter...
3.9K
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
569
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
569
Estimating Population Mean with Unknown Standard Deviation
8.0K
In practice, we rarely know the population standard deviation. In the past, when the sample size was large, this did not present a problem to statisticians. They used the sample standard deviation s as an estimate for σ and proceeded as before to calculate a confidence interval with close enough results. However, statisticians ran into problems when the sample size was small. A small sample size caused inaccuracies in the confidence interval.
William S. Gosset (1876–1937) of the...
William S. Gosset (1876–1937) of the...
8.0K
Distributions to Estimate Population Parameter
4.1K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.1K


