Related Experiment Video
Updated: Feb 12, 2026

Measuring Attention and Visual Processing Speed by Model-based Analysis of Temporal-order Judgments
Published on: January 23, 2017
A hybrid ResNet50-vision transformer model with an attention mechanism for aerial image classification
Amr Aboghanem1, Mohamed Abd Elfattah2, Hanan M Amer3
1Electronics and Communications Department, Faculty of Engineering, Mansoura University, Mansoura University, Mansoura, 35516, Egypt. Amraboghanem836@gmail.com.
None:
Aerial image classification is considered an open challenge due to its properties and the presence of various complex images. Given the complexity and variation in aerial images, this paper proposes two hybrid models for classification. The first hybrid model combines features extracted from ResNet-50 and the Vision Transformer (ViT), followed by the application of multi-head attention (MHA) to detect the most informative features. The second hybrid model also extracts features from ResNet-50 and ViT, then applies cross-attention. Both hybrid models are assessed using the benchmark Sikkim Aerial Images Dataset for Object Detection (SAIOD). The efficacy of the two hybrid models is assessed using the well-established performance metrics, including precision, recall, F1-score, and the ROC curve. The results indicate that the first model, which employs MHA, achieves superior performance with an accuracy of 95.80%. Both models outperform the best existing methods, achieving accuracies of 95.80% and 95.52%, respectively.
Related Concept Videos
Vision
The Quantum-Mechanical Model of an Atom
Color Vision
Hybrid Zones
Hybridization of Atomic Orbitals I
Bacterial Transformation
Griffith made an unexpected discovery when he killed the pathogenic strain and mixed its remains with the live, non-pathogenic strain. Not only did the mixture kill host mice, but it also contained living pathogenic bacteria that...

