Related Experiment Video
Updated: Sep 10, 2025

03:31
Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
635
Ensemble deep learning with image captioning for visual pollution detection, classification, and reporting
Haya Almalki1, Nahlah Algethami2
1Department of Computer Science, Saudi Electronic University, Riyadh, Saudi Arabia.
Scientific Reports
|August 26, 2025
Summary
This study introduces an AI framework using deep learning and image captioning to automatically detect and report visual pollution in Saudi cities. The system enhances urban management and reporting accuracy for sustainable development.
Area of Science:
- Environmental Science
- Computer Science
- Urban Planning
Background:
- Rapid urban development in Saudi Arabia, including Saudi Vision 2030 initiatives, has led to environmental challenges like visual pollution (VP).
- Current VP reporting methods rely on manual data entry via an online application, which is error-prone.
- There is a need for automated, accurate systems to monitor and manage urban environmental quality.
Purpose of the Study:
- To propose an AI-driven framework for automated detection, classification, and reporting of visual pollution.
- To integrate deep learning models (YOLOv5, EfficientDet) with ensemble techniques for enhanced VP detection.
- To utilize Bootstrapping Language-Image Pre-training (BLIP) for automatic generation of image-based report descriptions.
Main Methods:
- Developed an AI framework combining YOLOv5 and EfficientDet models for visual pollution detection.
- Employed ensemble techniques to improve the performance of individual deep learning models.
- Integrated BLIP-2 (specifically BLIP2-Flan-T5-XL) for automatic image captioning to generate report descriptions.
- Utilized the
- Saudi Arabia Public Roads Visual Pollution Dataset
- for framework development and evaluation.
Main Results:
- The ensemble approach achieved a Mean Average Precision (mAP) of 0.95, recall of 0.95, precision of 0.91, and F1 score of 0.93.
- Individual models were surpassed by the ensemble method in visual pollution detection.
- The BLIP2-Flan-T5-XL model demonstrated 80% accuracy in generating descriptive text for urban images based on human evaluation.
Conclusions:
- The proposed AI framework effectively automates visual pollution detection and reporting, significantly reducing errors associated with manual entry.
- The integration of deep learning and image captioning facilitates the creation of citizen-monitored reports, improving urban management.
- This AI-driven approach contributes to enhancing the quality of life and promoting sustainable urban development in Saudi cities.
Related Concept Videos
Aggregates Classification
381
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
381
Deconvolution
251
Deconvolution, also known as inverse filtering, is the process of extracting the impulse response from known input and output signals. This technique is vital in scenarios where the system's characteristics are unknown, and they must be inferred from the observable signals.
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
251

