结合卷积和视觉变压器的双分支模型用于作物疾病分类
Qingduan Meng1, Jiadong Guo1, Hui Zhang1
1College of Information Engineering, Henan University of Science and Technology, Luoyang, Henan, China.
PloS one
|April 24, 2025
概括
本研究引入了一种结合卷积神经网络 (CNN) 和视觉转换器 (ViT) 的新型双分支模型,用于准确的作物疾病分类. 轻量级模型在较少的参数下实现了高精度,为植物疾病识别提供了有效的解决方案.
科学领域:
- 农业科学 农业科学
- 计算机视觉 计算机视觉
- 机器学习 机器学习
背景情况:
- 农作物疾病分类对于粮食安全至关重要,但由于复杂的视觉特征而具有挑战性.
- 现有的计算机视觉模型经常与植物疾病的复杂纹理和形状作斗争.
- 卷积神经网络 (CNN) 和视觉转换器 (ViT) 的集成为改进分类提供了一个有希望的途径.
研究的目的:
- 开发一种轻量级且高度准确的双分支模型用于作物疾病分类.
- 有效地将本地特征提取 (CNN) 与全球特征理解 (ViT) 结合起来.
- 提高农业应用的变压器模型的表示能力.
主要方法:
- 一个双分支架构,将CNN集成为本地功能和ViT集成为全球功能.
- 引入聚合局部感知料前向层 (ALP-FFN) 以增强变压器的局部性.
- 开发一个轻量级的变压器块与ALP-FFN和线性自我注意力减少复杂性.
主要成果:
- 在仅有490万个参数的PlantVillage数据集上实现了99.71%的准确性,超过了最先进的模型.
- 在土豆叶数据集上获得了98.78%的准确性,超过了ResNet-18.
- 在显著降低参数和计算成本的情况下,证明了卓越的性能.
结论:
- 拟议的双分支CNN-ViT模型有效地利用本地和全球特征来精确识别作物疾病.
- 采用ALP-FFN的轻量级设计为现实世界农业应用提供了高效的解决方案.
- 这种方法为自动作物疾病诊断提供了一种强大而可扩展的方法.
相关概念视频
Light Acquisition
8.4K
In order to produce glucose, plants need to capture sufficient light energy. Many modern plants have evolved leaves specialized for light acquisition. Leaves can be only millimeters in width or tens of meters wide, depending on the environment. Due to competition for sunlight, evolution has driven the evolution of increasingly larger leaves and taller plants, to avoid shading by their neighbors with contaminant elaboration of root architecture and mechanisms to transport water and nutrients.
8.4K
Visual System
426
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
426
Vision
52.2K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
52.2K
Parallel Processing
125
The brain processes sensory information rapidly due to parallel processing, which involves sending data across multiple neural pathways at the same time. This method allows the brain to manage various sensory qualities, such as shapes, colors, movements, and locations, all concurrently. For instance, when observing a forest landscape, the brain simultaneously processes the movement of leaves, the shapes of trees, the depth between them, and the various shades of green. This enables a quick and...
125
Color Vision
370
Color perception begins in the retina, the light-sensitive layer at the back of the eye. Two main theories explain how colors are seen: the trichromatic theory and the opponent-process theory. The trichromatic theory, proposed by Thomas Young in 1802 and extended by Hermann von Helmholtz in 1852, suggests that color vision is based on three types of cone receptors in the retina. These cones are sensitive to different but overlapping ranges of wavelengths corresponding to red, blue, and green.
370


