机器学习算法的应用和特征选择在菜种子 (Brassica napus L.) 育种中用于种子产量
Masoud Shahsavari1, Valiollah Mohammadi2, Bahram Alizadeh3
1Department of Agronomy and Plant Breeding, College of Agriculture and Natural Resources, University of Tehran, Karaj, Iran.
Plant methods
|June 16, 2023
概括
机器学习通过识别关键特征,准确预测菜种子产量 (SY). 具有特征选择的多层感知神经网络为优化菜种子育种计划提供了一种有效的方法.
科学领域:
- 农业科学 农业科学
- 植物育种 植物育种
- 计算生物学 计算生物学
背景情况:
- 了解大麻种子产量 (SY) 和产量相关特征之间的关系对于高效的育种至关重要.
- 传统的方法难以处理SY和其他特征之间的复杂相互作用.
- 先进的机器学习 (ML) 是必要的,以有效的间接选择在菜种子的育种.
研究的目的:
- 确定ML算法和特征选择方法的最佳组合,以最大限度地提高RAPSEEDSY间接选择的效率.
- 探索各种ML模型对菜种子产量的预测能力.
主要方法:
- 采用了25个基于回归的ML算法和6种特征选择方法.
- 收集了两年 (2019-2021) 时间内从二十种菜种子基因型中收集的SY和产量相关数据.
- 使用根平均平方误差 (RMSE),平均绝对误差 (MAE) 和确定系数 (R2) 评估算法性能.
主要成果:
- 支持向量回归在所有15个特征中实现了最佳性能 (R2=0.860,RMSE=0.266,MAE=0.210).
- 多层感知神经网络 (MLPNN-Identity) 具有三个选定的特征,表现出高效率 (R2=0.843,RMSE=0.283,MAE=0.224).
- 对SY的关键预测特征包括每个植物的,生理成熟的几天,植物的高度和第一个的高度.
结论:
- MLPNN-Identity与逐步和倒向选择相结合,提供了一种可靠的方法,用于使用更少的特征准确预测SY.
- 这种基于机器学习的方法可以优化和加速油菜SY育种计划.
- 该研究强调了ML在提高作物改进间接选择效率方面的潜力.
相关概念视频
Plant Breeding and Biotechnology
19.5K
Crop cultivation has a long history in human civilization, with records showing the cultivation of cereal plants beginning at around 8000 BC. This early plant breeding was developed primarily to provide a steady supply of food.
19.5K
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K


