Related Experiment Video
Updated: May 2, 2026

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Addressing the data imbalance issue in machine learning modeling of rare and disruptive outage events
Morteza Azizi1,2, Xinxuan Zhang2,3, Tala Yasenpoor4
1School of Civil and Environmental Engineering, University of Connecticut, Storrs, CT, 06268, USA.
Abstract:
Rare natural hazards-such as severe storms, flooding, and wildfires-are difficult to model due to scarcity of high-quality observational data. This scarcity results in an imbalance within the dataset, where high-impact events are severely underrepresented, reducing the effectiveness of machine learning (ML) models. In this study, we propose a framework using stable diffusion to generate physical consistent synthetic storm events for data enrichment. Our method generates paired input-output samples, ensuring that synthetic meteorological fields are meaningfully aligned with their corresponding impact. A variational autoencoder compresses 19-channel storm fields into latent space, where a cluster-conditioned diffusion model generates events aligned with outage severity. Synthetic events are filtered using evaluation metrics that assess distributional similarity and physical consistency. Enrichment with screened synthetic events significantly improved ML-based outage prediction modeling, which consequently boosted overall prediction accuracy-reducing CRMSE by 39%, increasing R2 by 11%, and NSE by over 200%.
Related Concept Videos
Outliers and Influential Points
Detection of Gross Error: The Q Test
Quantifying and Rejecting Outliers: The Grubbs Test
Steps in Outbreak Investigation
Censoring Survival Data
Survival Tree
Building a Survival Tree
Constructing a...

