从数据到决策:使用可解释的人工智能来预测主要生产国的大豆产量
Xiangyi Wang1,2, Yingbin He3,4, Huicong Chen1,2
1State Key Laboratory of Efficient Utilization of Arid and Semi-arid Arable Land in Northern China, Institute of Agricultural Resources and Regional Planning, Chinese Academy of Agricultural Sciences, Beijing, China.
Scientific reports
|January 13, 2026
概括
像科尔莫戈罗夫-阿诺德网络 (KAN) 这样的可解释人工智能 (XAI) 模型提供了与深度学习 (DL) 模型可比的准确作物产量预测. KAN提高了透明度,这对于农业决策和全球粮食安全至关重要.
科学领域:
- 农业科学 农业科学
- 人工智能的人工智能
- 可解释的人工智能
背景情况:
- 准确的作物产量估计对于全球粮食安全和贸易至关重要.
- 深度学习 (DL) 模型在预测方面表现出色,但缺乏透明度,阻碍了信任.
- 可解释AI (XAI) 旨在平衡模型解释性与预测准确性.
研究的目的:
- 提出和评估XAI-Crop框架,以实现透明的作物产量估计.
- 为了比较Kolmogorov-Arnold网络 (KAN) 与多层感知器 (MLP) 和随机森林 (RF) 模型的性能.
- 通过使用多来源数据,评估收益率驱动因素的区域变异性.
主要方法:
- 开发了使用多源数据的XAI-Crop框架.
- 使用主要大豆生产国进行了比较实验.
- 评估了科尔莫戈罗夫-阿诺德网络 (KAN),多层感知器 (MLP) 和随机森林 (RF) 模型.
主要成果:
- 在小样本环境中,KAN表现出与MLP和RF相比的预测准确性和概括性.
- 与MLP和RF相比,KAN提供了更好的解释性.
- 太阳引起的叶绿素光 (SIF) 被确定为所有地区的一致敏感预测因素,突出了产量驱动因素的区域变化.
结论:
- 以KAN为例的XAI方法可以有效地弥合模型准确性和可解释性之间的差距.
- XAI-Crop框架显示了整合到农业决策支持系统的可行性.
- 通过提高产量估计的透明度,研究结果有助于可持续发展的农业发展.
相关概念视频
Multiple Regression
3.7K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.7K
Plant Breeding and Biotechnology
21.4K
Crop cultivation has a long history in human civilization, with records showing the cultivation of cereal plants beginning at around 8000 BC. This early plant breeding was developed primarily to provide a steady supply of food.
21.4K
What is Climate?
20.4K
Climate refers to the prevailing weather conditions in a specific area over an extended period. As the saying goes, “Climate is what you expect. Weather is what you get.” Climate is influenced by geographic factors, such as latitude, terrain, and proximity to bodies of water.
20.4K
Light Acquisition
9.4K
In order to produce glucose, plants need to capture sufficient light energy. Many modern plants have evolved leaves specialized for light acquisition. Leaves can be only millimeters in width or tens of meters wide, depending on the environment. Due to competition for sunlight, evolution has driven the evolution of increasingly larger leaves and taller plants, to avoid shading by their neighbors with contaminant elaboration of root architecture and mechanisms to transport water and nutrients.
9.4K
Prediction Intervals
3.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.3K
Regression Analysis
8.0K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
8.0K

