A statistical filtering procedure to improve the accuracy of estimating population parameters in feed composition
P S Yoder1, N R St-Pierre2, W P Weiss1
1Department of Animal Sciences, Ohio Agricultural Research and Development Center, The Ohio State University, Wooster 44691.
Journal of Dairy Science
|July 6, 2014
Summary
Accurate feed nutrient data is crucial for diet formulation. This study developed a mathematical filter to clean large feed composition databases, improving nutrient variance and correlation estimates for better risk assessment.
Area of Science:
- Animal Nutrition
- Data Science
- Agricultural Science
Background:
- Accurate nutrient composition, variance, and covariance are essential for quantitative diet formulation and risk reduction.
- Commercial feed databases may contain inaccurate data due to misidentified feeds or outliers, affecting statistical reliability.
Purpose of the Study:
- To design a mathematical filter for accurate estimation of nutrient distribution moments (mean, variance, covariance) within feed populations, even with outliers and subpopulations.
- To generate improved feed composition tables with precise means, variances, and correlations.
Main Methods:
- A data filtering procedure combining univariate, principal components analysis (PCA), and cluster analysis was applied to over 1.3 million feed samples.
- The procedure identified and removed outliers and distinct subpopulations within feed datasets.
Main Results:
- An average of 13.5% of samples were removed, with multivariate methods removing the majority.
- The filter significantly reduced standard deviations and altered correlation estimates, while means remained largely unaffected.
- Inaccurate feed identification and distinct subpopulations (e.g., immature vs. mature silage) were identified as sources of data contamination.
Conclusions:
- The developed procedure effectively refines feed composition data, providing more accurate estimates of nutrient variation and inter-nutrient relationships.
- Improved data accuracy enhances economic evaluations, risk assessments in diet formulation, and enables stochastic programming applications.
More Related Videos
Related Concept Videos
Mechanistic Models: Compartment Models in Individual and Population Analysis
359
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
359
Distributions to Estimate Population Parameter
4.5K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.5K
Contaminants and Errors
664
Effective sample preparation is crucial for accurate and reliable laboratory analysis. During this process, two significant sources of error can arise: concentration bias from improper sample splitting and contamination caused by methods used to reduce particle size, such as grinding or homogenization. Identifying and minimizing these potential errors is crucial to ensuring the validity of the analysis.
Another key consideration is determining the appropriate number of samples required to...
Another key consideration is determining the appropriate number of samples required to...
664
Analysis of Population Pharmacokinetic Data
992
Analysis of population pharmacokinetic data involves studying the behavior of drugs within diverse populations to understand their pharmacokinetic parameters. Traditional pharmacokinetic methods typically involve collecting samples from a few individuals and estimating these parameters. While these methods are commonly used, they have limitations in capturing the variability in drug response among individuals or heterogeneous populations. Population pharmacokinetics is employed to address these...
992
What are Estimates?
7.6K
It isn't easy to measure a parameter such as the mean height or the mean weight of a population. So, we draw samples from the population and calculate the mean height or mean weight of the individuals in the sample. This sample data acts as a representative measure of the population parameter. These sample statistics are known as estimates.
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...
7.6K
Estimating Population Mean with Known Standard Deviation
7.4K
To construct a confidence interval for a single unknown population mean μ, where the population standard deviation is known, we need sample mean as an estimate for μ and we need the margin of error. Here, the margin of error (EBM) is called the error bound for a population mean (abbreviated EBM). The sample mean is the point estimate of the unknown population mean μ.
The confidence interval estimate will have the form as follows:
(point estimate - error bound, point estimate +...
The confidence interval estimate will have the form as follows:
(point estimate - error bound, point estimate +...
7.4K


