Improved Random Forest Algorithm Based on Decision Paths for Fault Diagnosis of Chemical Process with Incomplete Data
Yuequn Zhang1, Lei Luo1, Xu Ji1
1Department of Chemical Engineering, Sichuan University, Chengdu 610065, China.
Sensors (Basel, Switzerland)
|October 26, 2021
Summary
This study introduces a novel fault detection and diagnosis method, DPRF, which accurately handles missing data using random forest decision paths and correction coefficients for improved industrial process monitoring.
Area of Science:
- Chemical Engineering
- Data Science
- Industrial Process Control
Background:
- Data-driven fault detection and diagnosis (FDD) is crucial for industrial processes.
- Existing FDD methods struggle with accuracy when data is missing.
- Big data analytics has increased the need for robust FDD techniques.
Purpose of the Study:
- To propose an improved random forest (RF) based FDD method, DPRF, that effectively compensates for incomplete data.
- To enhance the reliability of FDD in the presence of missing sensor readings.
- To demonstrate the superiority of DPRF over existing methods in handling data anomalies.
Main Methods:
- Developed a Decision Path Random Forest (DPRF) model incorporating correction coefficients.
- Utilized intact training samples to build decision trees within the RF.
- Inferred sample reliability scores from decision paths and node importance for each tree.
- Implemented a majority voting system combining predictions and reliability scores.
Main Results:
- The DPRF model demonstrated superior performance in fault detection and diagnosis with incomplete data.
- Tested on the Tennessee Eastman (TE) process, DPRF showed enhanced accuracy compared to other FDD methods.
- Correction coefficients effectively compensated for the influence of missing data points.
Conclusions:
- The proposed DPRF method offers a robust and accurate solution for FDD with missing data.
- DPRF significantly improves the reliability of fault diagnosis in industrial big data scenarios.
- This approach provides a valuable tool for maintaining operational integrity in complex industrial systems.
Related Concept Videos
Survival Tree
181
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
181
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
121
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
121
Decision Making: P-value Method
5.9K
The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
5.9K
Mechanistic Models: Compartment Models in Individual and Population Analysis
109
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
109
Propagation of Uncertainty from Random Error
1.3K
An experiment often consists of more than a single step. In this case, measurements at each step give rise to uncertainty. Because the measurements occur in successive steps, the uncertainty in one step necessarily contributes to that in the subsequent step. As we perform statistical analysis on these types of experiments, we must learn to account for the propagation of uncertainty from one step to the next. The propagation of uncertainty depends on the type of arithmetic operation performed on...
1.3K
Response Surface Methodology
319
Response Surface Methodology (RSM) is a collection of statistical and mathematical techniques used to develop, improve, and optimize processes. It is particularly valuable when many input variables or factors potentially influence a response variable.
The process of RSM involves several key steps:
The process of RSM involves several key steps:
319


