使用机器学习和部分依赖来评估表型值的最佳线性无偏预测 (BLUP) 的稳定性
Prashant Bhandari1, Tong Geon Lee2,3
1Horticultural Sciences Department, University of Florida, Gainesville, FL, 32611, USA.
Journal of applied genetics
|January 3, 2024
概括
植物研究中的最佳线性无偏预测 (BLUP) 对实验设计很敏感. 我们的研究表明,并非所有实验重复都对BLUP的准确性有同等的贡献,影响结果.
科学领域:
- 农业科学 农业科学
- 遗传学 是一个遗传学.
- 生物统计学 生物统计学
背景情况:
- 最好的线性无偏预测 (BLUP) 是植物研究中的标准方法,用于处理表型数据中的实验变异.
- BLUP的准确性严重依赖于受控的实验重复和可变组件的准确建模.
研究的目的:
- 通过评估单个实验重复的贡献来评估BLUP值的稳定性.
- 确定重复的不平等贡献如何影响表型数据集中的BLUP准确性.
主要方法:
- 利用机器学习技术,特别是神经网络,来估计实验重复的特征重要性.
- 使用部分依赖性分析计算单个重复对BLUP值的平均边际影响.
- 将神经网络的特征重要性与部分依赖性分析结果进行比较.
主要成果:
- 证明实验重复在表型数据集中对最终的BLUP值有不平等的贡献.
- 表明一些重复不成比例地影响BLUP结果,可能导致错过了真正的积极关联.
- 鉴定了不同实验重复对BLUP准确性影响的变异性.
结论:
- BLUP的稳定性受到实验重复的不平等贡献的影响.
- 在BLUP模型中概述变量组件的进一步细化是必要的,以解决重复的不成比例的影响.
- 改进BLUP模型规范可以提高植物研究的准确性和可靠性.
相关概念视频
Residuals and Least-Squares Property
7.4K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.4K
Variation
6.8K
An important characteristic of any set of data is the variation in the data. In some data sets, the data values are concentrated closely near the mean; in other data sets, the data values are more widely spread out from the mean. The most common measure of variation, or spread, is the standard deviation, which is the square root of variance.
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
6.8K
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Mechanistic Models: Compartment Models in Individual and Population Analysis
43
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
43
Biostatistics: Overview
246
Biostatistics plays a crucial role in understanding and analyzing data in healthcare and biology. Biostatisticians conduct experiments, gather evidence, and draw meaningful conclusions using statistical methods and techniques. Different variables form the foundation of biostatistical analysis, allowing researchers to understand and interpret data effectively. These variables are classified into different types, each serving a specific purpose in statistical analysis.
Discrete variables are...
Discrete variables are...
246


