Household electricity consumption prediction using database combinations, ensemble and hybrid modeling techniques
Gaikwad Sachin Ramnath1, R Harikrishnan2, S M Muyeen3
1Symbiosis Institute of Technology (SIT), Pune Campus, Symbiosis International (Deemed) University, Pune, India.
Abstract:
Household electricity consumption (HEC) is changing over time, depends on multiple factors, and leads to effects on the prediction accuracy of the model. The objective of this work is to propose a novel methodology for improving HEC prediction accuracy. This study uses two original datasets, namely questionnaire survey (QS) and monthly consumption (MC), which contain data from 225 consumers from Maharashtra, India. The original datasets are combined to create three additional datasets, namely QS + MC, QS equation (QsEq) + next month's consumptions, and QsEq + MC. Furthermore, the HEC prediction accuracy is boosted by applying different approaches, like correlation methods, feature engineering techniques, data quality assessment, heterogeneous ensemble prediction (HEP), and the hybrid model. Five HEP models are created using dataset combinations and machine learning algorithms. Based on the MC dataset, the random forest provides the best prediction of RMSE (36.18 kWh), MAE (25.73 kWh), and R2 (0.76). Similarly, QsEq + MC dataset adaptive boosting provides a better prediction of RMSE (36.77 kWh), MAE (26.18 kWh), and R2 (0.76). This prediction accuracy is further increased using the proposed hybrid model to RMSE (22.02 kWh), MAE (13.04 kWh), and R2 (0.92). This research work benefits researchers, policymakers, and utility companies in obtaining accurate prediction models and understanding HEC.
Related Concept Videos
Mechanistic Models: Compartment Models in Individual and Population Analysis
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Distribution Reliability and Automation
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Electrical Energy
Wind Turbine Machine Models
Induction machines interact through the rotating magnetic field generated by the stator and the rotor. The key parameter is slip, which is the difference between synchronous speed and rotor speed relative to synchronous speed. Slip is...


