Interpretable machine learning models to predict cadmium in wheat for safe production and soil management

Qi-Xin Lü1, Zhi-Xian Tang1, Zhong Tang1

  • 1Center for Agricultural and Environmental Health, Jiangsu Collaborative Innovation Center for Solid Organic Waste Resource Utilization, College of Resources and Environmental Sciences, Nanjing Agricultural University, Nanjing 210095, China.

Fundamental Research
|June 11, 2026
PubMed
Summary

Machine learning accurately predicts cadmium (Cd) in wheat grain using soil properties. The eXtreme Gradient Boosting (XGBoost) model identified soil Cd and pH as key factors, aiding food safety and sustainable agriculture.

Related Concept Videos

Multiple Regression01:25

Multiple Regression

Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Key Elements for Plant Nutrition02:35

Key Elements for Plant Nutrition

Like all living organisms, plants require organic and inorganic nutrients to survive, reproduce, grow and maintain homeostasis. To identify nutrients that are essential for plant functioning, researchers have leveraged a technique called hydroponics. In hydroponic culture systems, plants are grown—without soil—in water-based solutions containing nutrients. At least 17 nutrients have been identified as essential elements required by plants. Plants acquire these elements from the atmosphere, the...
Mechanistic Models: Compartment Models in Individual and Population Analysis01:23

Mechanistic Models: Compartment Models in Individual and Population Analysis

Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least squares (OLS)...