通过在阿拉比卡咖啡中堆叠合奏学习来提高基因组预测
Moyses Nascimento1,2, Ana Carolina Campana Nascimento1,2, Camila Ferreira Azevedo1
1Laboratory of Intelligence Computational and Statistical Learning (LICAE), Department of Statistics, Federal University of Viçosa, Viçosa, Brazil.
Frontiers in plant science
|August 1, 2024
概括
堆叠集体学习 (SEL) 显著提高了咖啡阿拉比卡繁殖中的基因组选择精度. 这种基于DNA的方法增强了对产量和抗病能力等关键特征的预测,优于传统方法.
科学领域:
- 植物育种与遗传学
- 生物信息学和计算生物学
- 农业科学 农业科学
背景情况:
- 传统的咖啡育种依赖于漫长的表型观测,阻碍了品种的快速发展.
- 基因组选择 (GS) 为识别优质咖啡 (Coffea Arabica) 基因型提供了一个更快,基于DNA的替代方案.
- 集体学习方法,特别是堆叠集体学习 (SEL),显示出在复杂特征选择中提高预测准确性的潜力.
研究的目的:
- 调查堆叠集体学习 (SEL) 的有效性,以提高咖啡阿拉比卡基因组选择中的预测准确性.
- 评估SEL在预测关键农业学和抗病特征方面的表现:产量 (YL),水果数量 (NF),叶矿工感染 (LM) 和纤维菌病发病率 (Cer).
主要方法:
- 分析了195个咖啡阿拉比卡个体的基因型,具有21,211个单核酸多态性 (SNP) 标记.
- 实施交叉验证 (CV) 方案来评估模型性能.
- 使用基因组最佳线性无偏预测 (GBLUP),MARS,QRF和RF作为SEL框架内的基础学习者,而Ridge回归,RF,GBLUP和单个平均值作为meta-learner.
主要成果:
- 堆叠集体学习 (SEL) 在与个人基础学习者模型相比,在所有评估的特征中表现出卓越的预测能力 (PA).
- 与GBLUP相比,SEL在PA方面取得了显著的收益:收益率为87.44% (YL),水果数为37.83% (NF),叶矿工感染为199.82% (LM),虫病发病率为14.59%.
- 该研究证实了SEL能够准确预测阿拉伯咖啡的重要特征的能力.
结论:
- 堆叠集体学习 (SEL) 是咖啡育种中基因组选择的一个非常有前途的进步.
- 通过整合来自多个模型的预测,SEL有效地提高了阿拉伯咖啡复杂特征的预测准确度.
- 这种方法加速了优质咖啡品种的选择,解决了传统育种方法的局限性.
更多相关视频
相关概念视频
Improving Translational Accuracy
9.7K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
9.7K
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Genomics
36.2K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
36.2K
Aggregates Classification
310
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
310
Extraction: Advanced Methods
436
Metal ions can be separated from one another by complexation with organic ligands–the chelating agent– to form uncharged chelates. Here, the chelating agent must contain hydrophobic groups and behave as a weak acid, losing a proton to bind with the metal. Since most organic ligands used in this process are insoluble or undergo oxidation in the aqueous phase, the chelating agent is initially added to the organic phase and extracted into the aqueous phase. The metal-ligand complex is...
436


