Related Experiment Video
Updated: Jan 6, 2026

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
Fusion of deep transfer learning models with Gannet optimisation algorithm for an advanced image captioning system
Tareq M Alkhaldi1, Mashael M Asiri2, Fahad Alzahrani3
1Department of Educational Technologies, Imam Abdulrahman bin Faisal University, Dammam, Saudi Arabia.
This study introduces an advanced image captioning system using deep learning and optimization algorithms to help visually impaired individuals understand images. The novel Fusion of Deep Transfer Learning Models and the Gannet Optimisation Algorithm for an Advanced Image Captioning System for Visual Disabilities (FDTLGO-AICSVD) model significantly improves descriptive accuracy.
Area of Science:
- Computer Vision
- Natural Language Processing
- Artificial Intelligence
Background:
- Automated image captioning is crucial for assisting visually impaired individuals by converting visual information into spoken or written descriptions.
- Existing methods face challenges in generating precise and context-aware captions, limiting their effectiveness for accessibility applications.
Purpose of the Study:
- To develop a novel Fusion of Deep Transfer Learning Models and the Gannet Optimisation Algorithm for an Advanced Image Captioning System for Visual Disabilities (FDTLGO-AICSVD).
- To enhance image captioning accuracy and descriptive quality for visually impaired users through precise image-to-text conversion.
Main Methods:
- Image preprocessing techniques including noise removal and contrast enhancement.
- Feature extraction using deep transfer learning models (DenseNet121, VGG19, MobileNetV2) and Term Frequency Inverse Document Frequency (TF-IDF).
- Hyperparameter optimization via the Gannet Optimization Algorithm (GOA) for precise caption generation.
Main Results:
- The FDTLGO-AICSVD model achieved a superior BLEU-4 score of 45.11% on the Flickr8k dataset and 58.91% on the Flickr30k dataset.
- Significantly higher CIDEr scores were recorded: 63.17 on Flickr8k and 69.81 on Flickr30k.
- Demonstrated enhanced descriptive accuracy and language generation capabilities across standard image captioning benchmarks.
Conclusions:
- The proposed FDTLGO-AICSVD model offers a robust and efficient solution for image captioning, particularly benefiting visually impaired individuals.
- The integration of deep transfer learning and the Gannet Optimization Algorithm leads to superior performance in generating accurate and context-aware image descriptions.
Related Concept Videos
Learning Disabilities
Dyslexia
Dyslexia is a...
Improving Translational Accuracy
Improving Translational Accuracy