SABO-ILSTSVR:一种基因组预测方法,基于改进的最小平方双支持向量回归
Rui Li1,2, Jing Gao1,2,3, Ganghui Zhou1,2
1College of Computer and Information Engineering, Inner Mongolia Agricultual University, Hohhot, China.
Frontiers in genetics
|July 1, 2024
概括
这项研究介绍了SABO-ILSTSVR,这是一种新的基因组预测模型,通过优化ILSTSVR方法来提高准确性. 它有效地解决了基因组选择中的过度拟合,以实现更快的作物育种.
科学领域:
- 植物育种 植物育种
- 基因组学就是基因组学.
- 机器学习在农业中的应用
背景情况:
- 基因组预测 (GP) 使用高密度单核酸多态 (SNP) 预测基因组估计繁殖值 (GEBV),加速作物改进.
- 过度装配是GP的常见挑战,因为SNP的数量高于样本,阻碍了准确的预测.
研究的目的:
- 开发一个优化的基因组预测模型,以克服基于标记物的繁殖过度匹配.
- 在现代植物育种计划中提高基因组选择的准确性和效率.
主要方法:
- 该研究提出了一种增强的最小平方双支持向量回归 (LSTSVR) 模型,称为ILSTSVR,通过整合拉索规范化术语.
- 开发了一种基于减去平均值的新型优化器 (SABO),可以自动调整ILSTSVR模型的参数,从而生成SABO-ILSTSVR模型.
- 使用四个不同的作物数据集来评估SABO-ILSTSVR模型的性能.
主要成果:
- 与现有广泛使用的基因组预测方法相比,SABO-ILSTSVR模型在所有测试的数据集中显示出优越或同等的性能.
- 优化方法有效地缓解了传统GP模型中固有的过拟合问题.
- 拟议的方法为植物育种中准确预测GEBV提供了强大的解决方案.
结论:
- 萨博-ILSTSVR模型代表了基因组预测的重大进步,为植物育种提供了更高的准确性和效率.
- 这种优化的方法为加速遗传收益和缩短繁殖周期提供了宝贵的工具.
- 该研究强调了将先进的机器学习技术整合到强大的基因组选择中的潜力.
相关概念视频
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K
Improving Translational Accuracy
10.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
10.0K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Comparing Copy Number Variations and SNPs
17.7K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.7K
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Regression Analysis
5.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.7K


