A comparison of decision tree-based algorithms for food discrimination using vibrational spectroscopy
Leandro P da Silva1, Micael D L Oliveira1, Javier E L Villa1
1Institute of Chemistry, University of Campinas (UNICAMP), Campinas 13081-970, SP, Brazil.
Food Chemistry
|May 29, 2025
Summary
Decision tree algorithms, including random forest, accurately identify food authenticity using spectroscopy. Random forest demonstrated superior performance in discriminating gluten-free bread and adulterated coconut water.
Area of Science:
- Analytical Chemistry
- Chemometrics
- Food Science
Background:
- Spectroscopic techniques like near-infrared (NIR) and Raman spectroscopy are valuable for food analysis.
- Machine learning algorithms offer powerful tools for interpreting complex spectroscopic data.
- Accurate food discrimination is crucial for quality control, authenticity verification, and consumer safety.
Purpose of the Study:
- To systematically evaluate decision tree-based algorithms (decision tree, random forest, XGBoost) for binary food discrimination.
- To assess the performance of these algorithms using NIR and Raman spectroscopy.
- To enhance the chemical interpretability of the models.
Main Methods:
- Near-infrared (NIR) and Raman spectroscopy were employed for data acquisition.
- Decision tree, random forest, and XGBoost algorithms were applied for classification tasks.
- Feature importance was assessed using traditional methods and an impurity reduction strategy.
- Optimization of algorithms involved evaluating figures of merit and splitting methods.
Main Results:
- Accuracy values up to 99% were achieved in discriminating gluten-containing vs. gluten-free bread and pure vs. sucrose-adulterated coconut water.
- Key spectral features identified include NIR bands related to water (OH) and protein (CH, NH), and Raman bands related to glycosidic C-O-C bonds.
- Random forest outperformed decision tree and partial least squares discriminant analysis in accuracy and provided more robust results than XGBoost.
- The impurity reduction strategy improved chemical interpretability.
Conclusions:
- Decision tree-based algorithms, particularly random forest, are highly effective for food discrimination using spectroscopic data.
- Spectroscopy combined with machine learning provides a powerful approach for ensuring food authenticity and quality.
- The developed methods offer enhanced chemical interpretability, aiding in the understanding of discriminatory spectral features.


