完善处罚回归:一种用于优化基因组预测中的规范化参数的新方法
Abelardo Montesinos-López1, Osval A Montesinos-López2, Federico Lecumberry3
1Centro Universitario de Ciencias Exactas e Ingenierías (CUCEI), Universidad de Guadalajara, Guadalajara 44430, Jalisco, México.
G3 (Bethesda, Md.)
|November 9, 2024
概括
一种调整Ridge回归的新方法提高了基因组预测的准确性. 这种方法通过优化处罚参数来增强植物育种中的育种价值估计,从而在预测性能方面取得了显著的收益.
科学领域:
- 遗传学 是一个遗传学.
- 量化遗传学 量化遗传学
- 生物信息学是一种生物信息学.
背景情况:
- 基因组选择 (GS) 对于有效估计育种价值至关重要.
- 回归是一种流行的基因组预测方法,但它的性能取决于最佳的惩罚参数调整.
研究的目的:
- 引入一种新的,更有效的方法来选择基因组预测的Ridge回归中的最佳惩罚参数.
- 通过使用现实世界数据集来评估拟议方法与传统方法的性能.
主要方法:
- 在Ridge回归中开发一种用于最佳惩罚参数选择的新型算法.
- 在14个真实植物育种数据集中对拟议方法和常规方法进行比较分析.
主要成果:
- 拟议的方法在14个数据集中的13个中超过了传统方法.
- 在数据集中观察到56.15%的预测准确度 (皮尔森相关性) 的显著增长.
- 在正常化平均平方误差方面没有发现显著的增长.
结论:
- 这种新的惩罚参数选择方法显示了在基因组预测中改善Ridge回归的巨大潜力.
- 采用这种方法可以提高植物育种计划中候选品种的选择.
- 这些发现支持高级参数调整的实用性,以实现更准确的基因组评估.
相关概念视频
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Improving Translational Accuracy
9.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
9.1K
Multiple Regression
2.9K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
2.9K
Regression Analysis
5.6K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.6K
Truncation in Survival Analysis
168
Truncation in survival analysis refers to the exclusion of individuals or events from the dataset based on specific criteria related to the time of the event. This exclusion can happen in two primary forms: left truncation and right truncation.
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are...
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are...
168
Genome-wide Association Studies-GWAS
12.5K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
12.5K


