不同的方法来估计脊回归BLUP中的收缩因子,用于基因组选择
Hamid Sahebalam1, Mohsen Gholizadeh2
1Department of Animal Science, Faculty of Animal and Aquatic Science, Sari Agricultural Sciences and Natural Resources University, Sari, Iran.
Scientific reports
|November 26, 2025
概括
在基因组选择中估计收缩因子 (λ) 的等位基因频率总和 (AF-RRBLUP) 方法提供了高预测准确性和低计算成本. 这种方法通常优于基因组选择的其他直接和间接方法.
科学领域:
- 定量遗传学 是一种定量遗传学.
- 动物繁殖 动物繁殖
- 统计基因组学 统计基因组学
背景情况:
- 收缩因子 (λ) 对于惩罚性回归方法至关重要,如基因组选择中的回归最佳线性无偏预测 (RRBLUP).
- λ调节标记效应估计,模型复杂性和预测准确性 (PA).
- 评估不同的 λ 估计方法对于优化基因组选择模型至关重要.
研究的目的:
- 评估RRBLUP中估计λ的八种方法,并将它们与贝叶斯C (BC) 进行比较.
- 评估这些方法在不同基因架构和标记密度的性能.
- 为基因组选择确定最有效的λ估计方法.
主要方法:
- 模拟基因组有6个染色体,每个染色体有100个QTL.
- 四种不同标记数 (3000或9000) 和遗传性 (h2 = 0.2或0.6) 的四种情景.
- 评估了八种RRBLUP λ估计方法 (MSE-RRBLUP,PCC-RRBLUP,AIC-RRBLUP,BIC-RRBLUP,DIC-RRBLUP,NM-RRBLUP,AF-RBLUP,RRBLUP-BC) 和贝叶斯C的方法,这些方法都得到了广泛的应用.
主要成果:
- 间接的λ估计方法通常比直接方法的预测准确度 (PA) 高.
- 基因频率总和 (AF-RRBLUP) 方法显示出高的PA和计算效率.
- 基于信息标准的方法 (AIC,BIC,DIC) 显示了最低的PA,与BayesC和AF-RBLUP相比观察到显著差异.
结论:
- 由于其高PA和低计算负担的平衡,AF-RRBLUP方法是基因组选择的推方法.
- 方法之间的PA差异实际上是显著的,强调了适当的λ估计的重要性.
- 进一步的研究可能会在不同的基因组架构和真实数据场景下探索AF-RRBLUP的稳定性.
相关概念视频
Genome-wide Association Studies-GWAS
15.2K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
15.2K
Truncation in Survival Analysis
558
Truncation in survival analysis refers to the exclusion of individuals or events from the dataset based on specific criteria related to the time of the event. This exclusion can happen in two primary forms: left truncation and right truncation.
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are...
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are...
558
Genetic Drift
42.8K
Natural selection—probably the most well-known evolutionary mechanism—increases the prevalence of traits that enhance survival and reproduction. However, evolution does not merely propagate favorable traits, nor does it always benefit populations.
42.8K
Residuals and Least-Squares Property
8.9K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
8.9K
Quantifying and Rejecting Outliers: The Grubbs Test
3.5K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
3.5K
Multiple Regression
3.7K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.7K


