Related Experiment Video
Updated: Mar 18, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Deep learning-driven image captioning: Progress through transformers and large language models
Priyanka Panchal1, Vishal Polara2, Siddaraj U3
1Department of Information Technology, Madhuben and Bhanubhai Patel Institute of Technology, The Charutar Vidya Mandal (CVM) University, New Vallabh Vidya Nagar, Gujarat, India.
This study introduces a new deep learning model for image captioning, outperforming existing methods with advanced vision transformers and LLMs. The novel cross-attention mechanism enhances visual-linguistic alignment for more human-like image descriptions.
Area of Science:
- Artificial Intelligence
- Computer Vision
- Natural Language Processing
Background:
- Traditional Convolutional Neural Network-Recurrent Neural Network (CNN-RNN) hybrids and existing transformer models have limitations in image captioning.
- Achieving robust multimodal alignment and enhancing caption diversity remain key challenges in the field.
Purpose of the Study:
- To propose a novel deep learning model for image captioning using an advanced vision transformer architecture and a powerful Large Language Model (LLM).
- To improve the alignment between linguistic context and visual features through a unique cross-attention mechanism.
Main Methods:
- Development of a novel deep learning architecture integrating a vision transformer with an LLM.
- Implementation of a unique cross-attention mechanism for deep alignment between visual and linguistic features.
- Extensive evaluation on benchmark datasets including MSCOCO, Flickr30K, and NoCaps.
Main Results:
- The proposed model demonstrates significant improvements over traditional and existing transformer-based approaches.
- Achieved state-of-the-art performance comparable to leading methods like GIT, BLIP-2, and CoCa on MSCOCO, Flickr30K, and NoCaps.
- Specific metrics on MS COCO include BLEU-4 (0.495), METEOR (0.390), and CIDEr (1.32).
Conclusions:
- The novel architecture sets a new performance benchmark for image captioning systems.
- The fusion strategy proves efficient, enabling more precise, contextually rich, and human-like image descriptions.
- This work advances multimodal AI systems, supporting Sustainable Development Goals 9 and 4.
Related Concept Videos
Transformers
The iron core has a substantial relative permeability. Therefore, the magnetic field lines generated due to the current in one winding are almost entirely confined within the core, such that the same magnetic flux permeates each turn of both...
Types Of Transformers
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
Improving Translational Accuracy
Improving Translational Accuracy
