Related Experiment Video
Updated: Jun 6, 2026

11:49
Cereal Crop Ear Counting in Field Conditions Using Zenithal RGB Images
Published on: February 2, 2019
Distilled vision transformers with CNN fusion for robust cashew apple maturity prediction
Sumalatha Lingamgunta1, Jeevaratnam Mudidana1, Durai Raj Vincent2
1Department of Computer Science and Engineering, University College of Engineering Kakinada, Jawaharlal Nehru Technological University Kakinada, Kakinada, Andhra Pradesh, India.
Frontiers in Plant Science
|May 7, 2026
Summary
This study introduces a lightweight vision transformer model for cashew apple maturity grading, achieving 90% accuracy. This approach enhances post-harvest management through efficient, data-driven ripeness assessment.
Area of Science:
- Agricultural Science
- Computer Vision
- Machine Learning
Background:
- Cashew apples are highly nutritious but have a short shelf life due to their delicate nature.
- Accurate maturity grading is crucial for optimizing cashew apple post-harvest handling, storage, and market value.
- Current grading methods may not be efficient or scalable for the agricultural industry.
Purpose of the Study:
- To develop an efficient and accurate system for cashew apple maturity grading using artificial intelligence.
- To leverage knowledge distillation and lightweight transformer models for improved performance with limited data.
- To create a computationally efficient model suitable for real-time agricultural applications.
Main Methods:
- A lightweight vision transformer (ViT) student model was trained using multi-granular knowledge distillation (KD) from a stronger teacher model.
- The distillation process incorporated soft-label supervision, attention transfer, and token-level feature regression.
- Auxiliary lightweight models (MobileNet, ConvNeXt, EdgeNeXt) were used in an ensemble approach for enhanced predictions.
Main Results:
- The proposed ensemble model (ViT-KD with EdgeNeXt) achieved 90% accuracy on the test split.
- Stratified fivefold cross-validation demonstrated a mean accuracy of 86.89% ± 2.89%, indicating stable generalization.
- The distilled ViT model achieved real-time inference at 8.79 ms per image, outperforming conventional CNN baselines.
Conclusions:
- Knowledge-distilled lightweight transformers are effective for data-efficient maturity grading of cashew apples.
- The developed system offers a computationally efficient and accurate solution for agricultural applications.
- This approach has the potential to significantly improve post-harvest management and reduce food waste.