Optimizing chlorophyll content prediction in tea leaves via spectral transformations and deep learning
Yuta Tsuchiya1, Keita Yoshida2, Yoshiki Ishiguro3
1Graduate School of Science and Technology, Shizuoka University, Shizuoka, Japan.
BMC Plant Biology
|December 1, 2025
Summary
Accurate chlorophyll estimation in tea leaves using spectral reflectance is crucial for precision agriculture. Self-Supervised Learning (SSL) with Standard Normal Variate (SNV) preprocessing yielded the best prediction results.
Area of Science:
- Plant physiology and spectroscopy
- Precision agriculture and remote sensing
- Machine learning for biochemical trait estimation
Background:
- Accurate chlorophyll estimation is vital for monitoring plant health and optimizing agricultural practices.
- Hyperspectral reflectance data offers a non-destructive method for assessing plant physiological status.
- Various preprocessing techniques and machine learning models can influence the accuracy of chlorophyll content prediction.
Purpose of the Study:
- To evaluate the impact of four preprocessing techniques (Original Reflectance, Continuum Removal, De-trending, Standard Normal Variate) on chlorophyll content prediction accuracy.
- To compare the performance of four machine learning models (1D-CNN, SSL, ViT, Conformer) for chlorophyll estimation in tea leaves.
- To determine the optimal pairing of preprocessing methods and model architectures for enhanced prediction accuracy.
Main Methods:
- Collected hyperspectral reflectance data from tea leaves (Camellia sinensis).
- Applied four preprocessing techniques: Original Reflectance (OR), Continuum Removal (CR), De-trending (DT), and Standard Normal Variate (SNV).
- Utilized four machine learning models: 1D Convolutional Neural Network (1D-CNN), Self-Supervised Learning (SSL), Vision Transformer (ViT), and Conformer, with ten-fold cross-validation.
Main Results:
- Standard Normal Variate (SNV) and De-trending (DT) preprocessing enhanced spectral sensitivity to chlorophyll, especially in key absorption regions.
- The Self-Supervised Learning (SSL) model combined with SNV preprocessing achieved the highest prediction accuracy (R² = 0.82, RPD = 2.37).
- Optimal preprocessing varied by model: 1D-CNN excelled with DT, while ViT and Conformer benefited most from CR.
Conclusions:
- The choice of preprocessing technique significantly impacts machine learning model performance for chlorophyll estimation.
- Model architecture and preprocessing method must be carefully paired to maximize prediction accuracy.
- Customized preprocessing strategies are essential for effective hyperspectral analysis and plant phenotyping.


