Related Experiment Video
Updated: Jun 25, 2025

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Enhancing credit scoring accuracy with a comprehensive evaluation of alternative data.
Rivalani Hlongwane1, Kutlwano K K M Ramaboa1, Wilson Mongwe2
1Graduate School of Business, University of Cape, Cape Town, South Africa.
This study shows that using alternative data, like social network defaults and regional economics, significantly improves credit scoring accuracy. These new predictors enhance predictive performance beyond traditional credit bureau data alone.
Area of Science:
- Financial Risk Management
- Data Science
- Machine Learning
Background:
- Traditional credit scoring models rely heavily on credit bureau data.
- Alternative data sources are often overlooked in credit risk assessment.
- Enhancing predictive accuracy in credit scoring remains a key challenge.
Purpose of the Study:
- To investigate the impact of alternative data on credit scoring model accuracy.
- To compare the performance of models using traditional versus combined data sources.
- To identify novel predictors for improved credit risk evaluation.
Main Methods:
- Analysis of a comprehensive home loan portfolio dataset from Home Credit Group.
- Application of the model-X knockoffs framework for systematic variable selection.
- Inclusion of alternative predictors such as social network default status, regional economic ratings, and local population characteristics.
Main Results:
- Credit scoring models incorporating alternative data demonstrated improved predictive performance.
- The enhanced models achieved an area under the curve (AUC) of 0.79360 on the Kaggle Home Credit default risk dataset.
- Performance surpassed models relying solely on traditional credit bureau data.
Conclusions:
- Leveraging diverse, non-traditional data sources significantly augments credit risk assessment.
- Alternative data enhances overall credit scoring model accuracy and predictive power.
- The study underscores the value of incorporating overlooked data for robust financial modeling.
More Related Videos
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Related Concept Videos
Detection of Gross Error: The Q Test
Quantifying and Rejecting Outliers: The Grubbs Test
Outliers and Influential Points
Reliability and Validity
Testing a Claim about Standard Deviation
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
Self-Evaluation: Self-Enhancement and Self-Verification