基于XGBoost的风险预测模型,使用消费者投诉进行大规模的汽车召回
Yi-Na Li1,2, Ming Jiang2, Likun Wang2
1School of Public Affairs, University of Science and Technology of China, Hefei, People's Republic of China.
概括
这项研究使用XGBoost来预测消费者投诉导致的车辆召回情况,为汽车制造商提供积极的风险管理. 这些模型准确预测回忆风险长达18个月,提高了安全性.
科学领域:
- 汽车工程 汽车工程
- 风险管理 风险管理
- 数据科学数据科学数据科学
背景情况:
- 车辆召回带来了重大的安全风险和经济负担.
- 由于复杂的数据和不断变化的因素,预测召回风险具有挑战性.
- 现有的方法往往忽略了结构化的投诉数据和不同的召回阶段.
研究的目的:
- 开发针对车辆召回风险的高精度预测模型.
- 通过消费者投诉,确定影响汽车召回的关键风险因素.
- 加强汽车制造商和监管机构的积极风险管理策略.
主要方法:
- 采用XGBoost机器学习模型进行深入分析.
- 利用了来自国家公路交通安全管理局 (NHTSA) 的综合数据.
- 综合结构化的消费者投诉数据和官方召回记录.
主要成果:
- 在各种时间窗口中实现了卓越的模型性能.
- 证明了高预测准确度和稳定性,曲线下面积值高达18个月.
- 区分初始和后续召回的指标,针对不同车辆生命周期阶段.
结论:
- XGBoost模型为积极的车辆召回风险管理提供了宝贵的支持.
- 结构化投诉数据的系统整合提高了回忆预测的准确性.
- 该研究弥合了在预测汽车生命周期不同阶段召回风险方面存在的关键差距.
相关概念视频
Quantifying and Rejecting Outliers: The Grubbs Test
3.6K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
3.6K
Prediction Intervals
3.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.3K
Multiple Regression
3.8K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.8K
Regression Analysis
8.1K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
8.1K
Relative Risk
2.0K
Relative risk (RR) is a statistical measure commonly used in epidemiology to compare the likelihood of a particular event occurring between two groups. This metric is important for evaluating the relationship between exposure to a specific risk factor and the probability of a particular outcome. It plays a crucial role in medical research, public health studies, and risk assessment. Relative risk quantifies how much more (or less) likely an event is to occur in an exposed group compared to an...
2.0K
Aggregates Classification
972
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
972

