Related Experiment Video
Updated: Mar 29, 2026

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Risk-Sensitive Machine Learning for Financial Decision Modeling Under Imbalanced Data: Evidence from Bank
Bowen Dong1, Xinyu Zhang2, Yang Liu3
1School of Electrical Automation and Information Engineering, Tianjin University, Tianjin 300072, China.
None:
Bank telemarketing campaigns often experience low subscription rates due to customer heterogeneity and severe class imbalance, which pose challenges for reliable predictive modeling. This study investigates a data-driven approach that integrates synthetic minority oversampling and cost-sensitive learning to improve the prediction of telemarketing outcomes. Experiments are conducted using the Portuguese Bank Marketing dataset, comprising 41,188 instances with a positive response rate of 11.3%. Eight machine learning models are evaluated under a unified preprocessing pipeline and five-fold stratified cross-validation, including Logistic Regression, Decision Tree, Random Forest, and Ensemble methods. The results show that Ensemble models, particularly CatBoost, XGBoost, and LightGBM, achieve improved performance compared with traditional baselines, with notable gains in minority-class recall and overall discrimination ability. The best-performing model attains an F1-score of 0.540, a recall of 0.812 for the positive class, and a ROC-AUC of 0.908. To enhance interpretability, SHAP-based analysis is applied to quantify feature contributions, identifying campaign duration, previous contact outcomes, and selected macroeconomic indicators as key predictors. These findings indicate that combining resampling strategies with cost-sensitive optimization provides a robust and transparent approach for learning from imbalanced telemarketing data, thereby supporting reproducible and data-driven financial decision-making by explicitly addressing difficulty in minority-class identification under imbalance and class imbalance under cross-entropy training in imbalanced banking data.
Related Concept Videos
Survival Tree
Building a Survival Tree
Constructing a...
Mathematical Modeling: Problem Solving
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as: