相关实验视频
可解释的人工智能驱动的客户流失预测:采用基于SHAP的特征分析多模型组合方法
Ali El Attar1, Mohammed El-Hajj1
1Faculty of Computer Studies (FCS), Arab Open University (AOU), Beirut, Lebanon.
Frontiers in artificial intelligence
|February 26, 2026
概括
这项研究利用机器学习增强了电信领域的客户流失预测. 梯度增强模型,特别是XGBoost,表现出卓越的表现,确定合同类型和任期作为针对性保留策略的关键流失指标.
科学领域:
- 电信 电信服务 电信服务 电信服务
- 机器学习 机器学习
- 数据科学数据科学数据科学
背景情况:
- 客户流失对电信的利能力构成重大挑战.
- 有效的保留策略对于保持市场份额和收入至关重要.
研究的目的:
- 开发和评估用于准确预测客户流失的机器学习框架.
- 确定电信部门客户流失的关键驱动因素.
- 为优化客户保留策略提供可操作的见解.
主要方法:
- 使用了Telco客户流失数据集 (n = 7,043).
- 实现特征工程和SMOTE过量采样.
- 训练和评估了七个机器学习模型,包括XGBoost,随机森林和多层感知器.
- 采用SHAP分析用于模型解释和客户细分.
主要成果:
- 梯度增强算法 (XGBoost,LightGBM,梯度增强) 实现了最高的平衡性能 (F1得分:0.84).
- XGBoost表现出优越的区分能力 (AUC-ROC:0.932).
- 合同类型,任期和技术支持被确定为主要的流失预测因素.
- 优化预测值 (0.528) 平衡精度 (0.90) 和回忆 (0.91),减少了15%的错误负数.
结论:
- 机器学习,特别是梯度增强,为电信中客户流失预测提供了一个强大的方法.
- 识别关键的流失驱动因素可以开发有针对性和有效的客户保留计划.
- 这些发现支持数据驱动的决策,用于电信客户保留努力.
相关概念视频
Prediction Intervals
3.5K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.5K
Multiple Regression
4.2K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
4.2K