Related Experiment Videos
Calibrated, cost-sensitive, explainable ensemble learning for customer churn prediction in retail banking: a
Farhana Karim1, Md Ishtiaque Alam2, Aram Lalafaryan3
1Harrison College of Business and Computing, Southeast Missouri State University, Cape Girardeau, MO, United States.
Introduction:
Customer churn prediction in retail banking poses significant financial challenges, demanding models that are both accurate and interpretable under regulatory constraints. Existing approaches often suffer from data leakage, poor probability calibration, and inadequate handling of class imbalance.
Methods:
We present a calibrated, cost-sensitive, and explainable ensemble learning pipeline evaluated on the publicly available Bank Customer Churn dataset (ChurnModelling.csv, n = 10, 000). Seven ensemble classifiers - Random Forest, Extra-Trees, XGBoost, LightGBM, CatBoost, Stacking, and Voting - alongside a Logistic Regression baseline were compared using a strictly leakage-safe design incorporating class-weighting, decision-threshold optimization, isotonic probability calibration, Bayesian hyperparameter optimization via Optuna, and repeated stratified cross-validation with Friedman and Nemenyi post hoc testing.
Results:
The Bayesian-tuned LightGBM achieved the highest CV ROC-AUC of 0.8649 ± 0.0096 and test ROC-AUC of 0.8668, PR-AUC of 0.7139, and Brier score of 0.1395. At the cost-sensitive threshold (0.265), recall reached 0.7592 with precision of 0.5091. Cumulative gains analysis showed that targeting the top 40% of customers by predicted churn probability captures approximately 80% of actual churners.
Discussion:
The pipeline demonstrates that class-weighting combined with threshold optimization is a calibration-preserving alternative to SMOTE oversampling. TreeSHAP analysis identifies Age, NumOfProducts, and IsActiveMember as the most influential predictors. Limitations include reliance on a single static dataset and the exploratory nature of the statistical significance claims. The framework is presented as a reproducible, leakage-safe single-dataset case study rather than a validated general solution.