Related Experiment Video
Updated: Feb 8, 2026

Methods of Soil Resampling to Monitor Changes in the Chemical Concentrations of Forest Soils
Published on: November 25, 2016
Predicting reference soil groups using legacy data: A data pruning and Random Forest approach for tropical
Kpade O L Hounkpatin1, Karsten Schmidt2, Felix Stumpf3
1University of Bonn, Institute of Crop Science and Resource Conservation (INRES), Soil Science and Soil Ecology, Nussallee 13, D-53115, Bonn, Germany. hozias@uni-bonn.de.
Data pruning improves soil group prediction accuracy. Using Random Forest with a 90% core dataset and feature elimination enhanced predictions for minor soil groups, showing substantial agreement.
Area of Science:
- Soil Science
- Machine Learning
- Data Science
Background:
- Predicting soil taxonomic classes is difficult due to data irregularities from multiple surveyors.
- Source errors in soil datasets can significantly impact classification accuracy.
Purpose of the Study:
- To investigate if data pruning methods improve the prediction accuracy of minor soil groups.
- To compare data pruning with random oversampling for Random Forest classification.
Main Methods:
- Applied data pruning to create subsets of Plinthosols data (80% and 90% core range).
- Utilized Random Forest (RF) with recursive feature elimination.
- Compared pruned datasets against the entire dataset and random oversampling.
Main Results:
- The best prediction accuracy was achieved using RF with recursive feature elimination on the 90% core range, non-oversampled dataset.
- This model showed substantial agreement (kappa=0.57) and increased prediction accuracy for smaller soil groups by 7% to 35%.
- Wetness index was identified as the primary factor influencing soil groups in the Dano catchment.
Conclusions:
- Data pruning, specifically using a 90% core range dataset, effectively reduces source errors and enhances soil group prediction.
- Random Forest with recursive feature elimination is a robust method for improving classification accuracy in imbalanced soil datasets.
- Soil moisture distribution, indicated by the wetness index, is a key determinant of reference soil group distribution.
More Related Videos
Related Concept Videos
Data Reporting and Recording
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
How Data are Classified: Numerical Data
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Model Approaches for Pharmacokinetic Data: Compartment Models
Two primary types of compartment models are recognized: mammillary and catenary. The more...
Model Approaches for Pharmacokinetic Data: Physiological Models
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...

