Related Experiment Videos
Vision-Based Environmental Sensing for Flood Risk Forecasting: Dataset Relabeling and Temporal Multi-Task Learning.
1Department of Smart ICT Convergence Engineering, Seoul National University of Science and Technology, Seoul 01811, Republic of Korea.
Sensors (Basel, Switzerland)
|June 12, 2026
Summary
This study reformulates CCTV flood data and uses temporal modeling for better flood risk forecasting. An image-only model outperformed multimodal approaches, highlighting the need for data consistency in flood prediction systems.
Area of Science:
- Environmental Science
- Computer Science
- Artificial Intelligence
Background:
- River flooding and urban inundation necessitate advanced forecasting systems capable of predicting future risks.
- Existing closed-circuit television (CCTV)-based flood datasets often suffer from imbalanced or temporally inconsistent risk labels.
- Current image-based flood analysis methods are largely confined to static scene understanding.
Purpose of the Study:
- To propose a dataset reformulation and temporal multi-task forecasting framework for CCTV-based flood-risk prediction.
- To address limitations in existing flood datasets and image-based approaches for dynamic risk assessment.
- To investigate the effectiveness of multimodal sensor fusion in flood forecasting.
Main Methods:
- A site-relative relabeling strategy was developed to convert noisy frame-level annotations into four distinct risk levels using visual and environmental cues.
- The dataset was transformed from frame-based to site-hour sequences to enable multi-horizon forecasting (1, 3, and 6 hours).
- Image-only, weather-only, and naive multimodal configurations were evaluated to assess sensor fusion limitations.
Main Results:
- The reformulated dataset enabled an image-only temporal model to achieve superior performance, with a mean Intersection over Union (mIoU) of 0.892 and a Dice score of 0.940.
- Naive multimodal fusion significantly degraded performance, reducing macro-averaged F1 score (Macro-F1) to 0.267 and high-risk recall to 0.070.
- Ablation studies revealed that temporal modeling was crucial, with its removal decreasing Macro-F1 to 0.227 and high-risk recall to 0.000.
Conclusions:
- Dataset reformulation and temporal modeling are essential for advancing CCTV-based flood analysis from static estimation to dynamic risk forecasting.
- The study underscores the challenges of multimodal sensor fusion when dealing with noisy, weakly correlated, or temporally misaligned cross-modal signals.
- Robust cross-modal alignment is a prerequisite for achieving reliable performance gains through multimodal sensing in flood prediction.