改进了基因组预测性能,使用各种模型的合奏
Shunichiro Tomura1,2, Melanie J Wilkinson1,2, Mark Cooper1,2
1Queensland Alliance for Agriculture and Food Innovation (QAAFI), Centre for Crop Science, The University of Queensland, St Lucia, QLD 4072, Australia.
G3 (Bethesda, Md.)
|March 4, 2025
概括
将多个基因组预测模型组合到一个合奏中,可以提高作物育种的选择精度. 这种整体方法提高了预测准确度,并减少了与单个模型相比的错误,加速了遗传收益.
科学领域:
- 植物育种 植物育种
- 基因组学就是基因组学.
- 统计遗传学 统计遗传学
背景情况:
- 基因组预测模型旨在提高选择精度,以在作物育种中获得更快的遗传收益.
- "没有免费午餐"定理表明,在所有场景中,没有单一的基因组预测模型是普遍优越的.
- 个别模型的局限性需要探索像模型集团这样的替代方法.
研究的目的:
- 为了研究将多个基因组预测模型结合成一个合奏的有效性.
- 为了利用多样性预测定理,它假定集体预测误差低于单个模型误差.
- 为了提高基因组预测的准确性和加速作物育种计划中的遗传收益.
主要方法:
- 开发了一种"天真"的整体平均模型,该模型对个体模型的预测具有同等权重.
- 通过使用两种与作物产量相关的特征 (日至合成,每种植物的机数) 评估了整体模型.
- 使用teosinte嵌套关联映射数据集进行模型评估.
主要成果:
- 与个人基因组预测模型相比,整体方法显示了预测准确度的提高.
- 整体建模导致研究特征的预测错误减少.
- 整体的好处归因于个别模型中预测的多样性.
结论:
- 整体基因组预测模型比单个模型提供了更好的性能.
- 整体方法能够更全面地理解复杂的特征基因组架构.
- 集体方法具有显著的潜力,可以加速作物育种计划中的遗传收益.
相关概念视频
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K
Improving Translational Accuracy
8.5K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
8.5K
Sensitivity, Specificity, and Predicted Value
158
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
158
Genome-wide Association Studies-GWAS
12.3K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
12.3K
Multiple Regression
2.9K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
2.9K
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K


