Related Experiment Video
Updated: Jul 12, 2026

Micron-scale Phenotyping Techniques of Maize Vascular Bundles Based on X-ray Microcomputed Tomography
Published on: October 9, 2018
MHENet: a multimodal hybrid enhancement network for accurate classification of unsound maize kernels
Chenxia Wan1,2, Ruitao Li1,2, Wenzheng Li1,2
1Key Laboratory of Grain Information Processing and Control (Henan University of Technology), Ministry of Education, Zhengzhou, China.
Abstract:
Accurate classification of unsound maize kernels is critical for grain quality detection and grading. Single red-green-blue (RGB) images are incapable of effectively capturing variations in the internal chemical components of maize kernels, which easily causes category confusion. Meanwhile, hyperspectral images (HSI) lacks the spatial structural information and exhibits the limited discriminative ability when used independently. To address these challenges, this paper proposes a multimodal hybrid enhancement network (MHENet) with a parallel spectral-spatial dual-branch architecture. The proposed network adopts an optimized 1DCNN-3 module for hyperspectral sequence feature extraction and MobileNetV4-Small for lightweight RGB image spatial feature extraction, enabling the comprehensive capture of both spectral physicochemical characteristics and external morphological appearance features of maize kernels. Specifically, the cross-modal projection embedding (CMPE) module achieves heterogeneous feature space alignment; the bidirectional Mamba attention (BMA) module extracts global contextual information with linear complexity; and the entropy-driven adaptive gating (EAG) fusion module enables dynamic weighted fusion of multimodal features. Experiments are conducted on a self-constructed RGB and HSI dataset of unsound maize kernels. The experimental results demonstrate that the proposed MHENet achieves an accuracy of 97.68%, a macro-P of 97.61%, a macro-R of 97.28%, and a macro-F1 of 97.44% in six-category maize kernel classification. Five-fold cross-validation results in an average accuracy of 97.62% and a standard deviation of 0.14. Compared with feature concatenation, Hadamard product, and cross-attention fusion strategies, the proposed MHENet outperforms these methods in terms of accuracy, lightweight, and stability. This work provides an essential guidance to accurately classify unsound maize kernels.
