Developing machine learning models for relative humidity prediction in air-based energy systems and environmental
Kinza Qadeer1, Ashfaq Ahmad2, Muhammad Abdul Qyyum1
1School of Chemical Engineering, Yeungnam University, Gyeongsan, 712-749, Republic of Korea.
Journal of Environmental Management
|May 16, 2021
Summary
Machine learning, specifically the random forest algorithm, accurately predicts relative humidity using temperature data. This approach offers a significant improvement over traditional methods for environmental management and energy system design.
Area of Science:
- Environmental Science
- Data Science
- Machine Learning
Background:
- Relative humidity prediction is complex due to its nonlinear behavior.
- Machine learning offers powerful solutions for nonlinear and complex problems.
- The random forest algorithm is effective, requiring minimal preprocessing.
Purpose of the Study:
- To implement the random forest approach for relative humidity estimation.
- To evaluate the model's performance against varying wet-bulb depressions.
- To compare the random forest model with a support vector regression model.
Main Methods:
- Utilized the random forest algorithm to estimate relative humidity based on dry- and wet-bulb temperatures.
- Integrated Aspen HYSYS V10 with MATLAB 2019a for a data mining environment.
- Assessed model robustness using varying wet-bulb depressions and compared with support vector regression.
Main Results:
- The random forest model achieved a 1.1% mean absolute deviation in relative humidity prediction compared to Aspen HYSYS.
- Model performance decreased with higher wet-bulb depressions (around 20.0 °C).
- The random forest model outperformed the support vector regression model by 74.4%.
Conclusions:
- The random forest approach provides a robust and accurate method for relative humidity estimation.
- This technique offers significant benefits for designing air-dependent energy systems and environmental management.
- Further research may explore model performance at higher wet-bulb depressions.
Keywords:
Air quality parametersAspen hysysEnvironmental management operationsMachine learning-based estimationRandom forestSupport vector machineMore Related Videos
Related Concept Videos
Regression Analysis
6.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
6.7K
Heating and Cooling Curves
25.1K
When a substance—isolated from its environment—is subjected to heat changes, corresponding changes in temperature and phase of the substance is observed; this is graphically represented by heating and cooling curves.
For instance, the addition of heat raises the temperature of a solid; the amount of heat absorbed depends on the heat capacity of the solid (q = mcsolidΔT). According to thermochemistry, the relation between the amount of heat absorbed or released by a substance, q, and its...
For instance, the addition of heat raises the temperature of a solid; the amount of heat absorbed depends on the heat capacity of the solid (q = mcsolidΔT). According to thermochemistry, the relation between the amount of heat absorbed or released by a substance, q, and its...
25.1K
Multiple Regression
3.4K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.4K
Classification of Systems-I
384
Linearity is a system property characterized by a direct input-output relationship, combining homogeneity and additivity.
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
384
Precipitation Processes
2.0K
The experimental conditions in a gravimetric analysis should be optimized to maximize the particle size and purity of the obtained precipitate. Ideally, the concentration of the precipitating reagent should be low with effective stirring to maintain low relative supersaturation for the growth of large crystals. In homogeneous precipitation, the precipitant is slowly generated by a chemical reaction in the solution to avoid local reagent excesses. For example, urea decomposes gradually to...
2.0K
Classification of Systems-II
295
Continuous-time systems have continuous input and output signals, with time measured continuously. These systems are generally defined by differential or algebraic equations. For instance, in an RC circuit, the relationship between input and output voltage is expressed through a differential equation derived from Ohm's law and the capacitor relation,
295


