Enhancing customer retention in telecom industry with machine learning driven churn prediction
Alisha Sikri1, Roshan Jameel2, Sheikh Mohammad Idrees3
1Noida Institute of Engineering and Technology, Greater Noida, 201306, Uttar Pradesh, India.
Scientific Reports
|June 7, 2024
Summary
Predicting customer churn is vital for business retention. A novel Ratio-based data balancing technique significantly improves machine learning model accuracy for identifying potential churners.
Area of Science:
- Machine Learning
- Data Science
- Business Analytics
Background:
- Customer churn poses a significant business challenge, necessitating effective prediction for retention strategies.
- Imbalanced and diverse customer data distributions complicate accurate churn prediction models.
- Existing literature highlights the need for advanced data balancing techniques in churn analysis.
Purpose of the Study:
- To introduce and evaluate a novel Ratio-based data balancing technique for improving churn prediction accuracy.
- To compare the effectiveness of the proposed technique against traditional data resampling methods.
- To assess the performance of various machine learning algorithms, including ensemble methods, on balanced datasets.
Main Methods:
- Development of a novel Ratio-based data balancing technique to address data skewness.
- Evaluation of machine learning algorithms: Perceptron, Multi-Layer Perceptron, Naive Bayes, Logistic Regression, K-Nearest Neighbour, Decision Tree, Gradient Boosting, and Extreme Gradient Boosting (XGBoost).
- Comparison of the proposed Ratio-based technique with traditional Over-Sampling and Under-Sampling methods using metrics like Accuracy, Precision, Recall, and F-Score.
Main Results:
- The Ratio-based data balancing technique demonstrated superior performance over traditional methods in churn prediction.
- Ensemble algorithms, specifically Gradient Boosting and XGBoost, outperformed single machine learning models.
- The XGBoost method with a 75:25 ratio yielded the most promising results for churn prediction accuracy.
Conclusions:
- The Ratio-based data balancing technique is effective in enhancing churn prediction accuracy by addressing data imbalance.
- Ensemble machine learning methods, particularly XGBoost, are highly effective for customer churn prediction when applied to balanced datasets.
- Optimized data balancing and ensemble modeling offer significant improvements for customer retention strategies.
Related Concept Videos
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Reducing Line Loss
150
In a three-phase circuit, line loss is an indicator of energy dissipated as heat due to the resistance of transmission lines. To address this, incorporating transformers into the system—a step-up transformer at the source and a step-down transformer at the load—is a strategic solution. Two three-phase transformers are introduced to improve this.
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss...
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss...
150
Regression Analysis
5.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.7K
Classification of Signals
441
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
441
End Point Prediction: Gran Plot
314
A Gran plot is used to predict the equivalence volume or endpoint of a potentiometric or acid-base titration without reaching the endpoint. Typically, titration data is collected as a function of the titrant's volume up to a point less than the equivalence volume and then transformed into a linear format. The straight line is extended to the x-axis, indicating the necessary titrant volume to achieve the equivalence point.
For potentiometric titration, the Gran plot is created by plotting...
For potentiometric titration, the Gran plot is created by plotting...
314
Distribution Reliability and Automation
107
Distribution reliability in electrical power systems is critical for ensuring an uninterrupted power supply to consumers at minimal cost. According to IEEE Standard Terms, reliability is the probability that a device will function without failure over a specified time period or amount of usage. For electric power distribution, this translates to maintaining continuous power supply and addressing customer concerns over power outages. Several indices, as defined by IEEE Standard 1366-2012, are...
107


