Related Experiment Video
Updated: Jan 13, 2026

05:37
An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
2.5K
Credit risk prediction model for listed companies based on improved reinforcement learning and Bayesian optimization
Cai Yuanqing1, Zhenming Gao2, Zhang Jian3
1College of Business and Public Management, Wenzhou-Kean University, Wenzhou, China.
Plos One
|October 28, 2025
Summary
This study introduces a novel credit risk prediction method using reinforcement learning (Proximal Policy Optimization) for feature selection and imbalanced classification. The approach significantly improves prediction accuracy for listed companies.
Area of Science:
- Financial Risk Management
- Machine Learning in Finance
- Computational Economics
Background:
- Credit risk prediction is crucial for financial institutions, regulators, and investors due to rapid financial sector growth.
- Traditional credit risk models (z-score, logit, KVM, neural networks) face challenges in feature selection, imbalanced classification, and hyperparameter optimization.
- Existing methods often yield suboptimal results, necessitating advanced approaches.
Purpose of the Study:
- To develop an advanced credit risk prediction model for publicly traded companies.
- To address limitations of traditional models by integrating reinforcement learning and advanced optimization techniques.
- To enhance the accuracy and efficiency of credit risk forecasting.
Main Methods:
- Utilized an off-policy Proximal Policy Optimization (PPO) algorithm, a reinforcement learning (RL) technique, for effective feature selection and imbalanced classification.
- Employed Bayesian Optimization Hyperband (BOHB) for efficient hyperparameter optimization, merging Bayesian optimization with Hyperband.
- Validated the model on diverse datasets: CSMAR, MorningStar, KMV default, GMSC, and UCICCD.
Main Results:
- The proposed model achieved superior performance compared to state-of-the-art methods across all tested datasets.
- Specific F-measures obtained were 90.763% (CSMAR), 86.358% (MorningStar), 87.047% (KMV default), 90.576% (GMSC), and 89.485% (UCICCD).
- Demonstrated enhanced sample efficiency and optimized data utilization through RL-based PPO.
Conclusions:
- The novel credit risk prediction method shows significant promise and efficiency in financial settings.
- The integration of PPO and BOHB represents a major advancement in credit risk assessment systems.
- The findings support the model's effectiveness in enhancing investigative approaches for credit risk.
Related Concept Videos
Prediction Intervals
3.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.2K
Equity Theory
252
Equity theory explains how our sense of fairness influences the dynamics of close relationships. Rooted in social psychology, the theory posits that individuals evaluate fairness by comparing the ratio of their contributions to the rewards they receive. Relationship satisfaction is highest when these ratios are perceived as balanced between partners, promoting mutual reciprocity and a sense of justice.Equity vs. Equality in RelationshipsEquity is distinct from equality. Fairness does not...
252
Reinforcement
816
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
816
Observational Learning
817
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
817
Actuarial Approach
284
The actuarial approach, a statistical method originally developed for life insurance risk assessment, is widely used to calculate survival rates in clinical and population studies. This method accounts for participants lost to follow-up or those who die from causes unrelated to the study, ensuring a more accurate representation of survival probabilities.
Consider the example of a high-risk surgical procedure with significant early-stage mortality. A two-year clinical study is conducted,...
Consider the example of a high-risk surgical procedure with significant early-stage mortality. A two-year clinical study is conducted,...
284