在企业信用风险预测中优化和应用SMOTE算法,同时考虑多元化战略
1Faculty of Economics and Management, Xi'an Kedagaoxin University, Xi'an, 710000, Shaanxi, China. weih53476@outlook.com.
Scientific reports
|July 2, 2025
概括
这项研究优化了合成少数人过量采样技术 (SMOTE) 用于企业信用风险预测,提高了超过21%的准确性,并加强了多元化公司的金融风险管理.
科学领域:
- 金融风险管理 金融风险管理
- 金融中的机器学习
- 企业金融公司财务
背景情况:
- 企业多元化使信用风险预测变得复杂,原因是各业务部门的财务结构和业绩异质.
- 金融机构需要对多元化企业的跨部门风险进行先进的监测和评估.
- 现有的信用风险模型面临的挑战是公司金融中常见的不平衡数据集.
研究的目的:
- 优化合成少数群体过量采样技术 (SMOTE) 算法,以提高企业信用风险预测.
- 分析企业多元化战略对信用风险评估的影响.
- 为金融机构提供改进的算法工具,用于在复杂的企业环境中管理风险.
主要方法:
- 开发了一个优化的SMOTE算法,包括一个自适应边界调整机制和一个优化的重量分配协议.
- 系统地分析公司多元化战略及其财务影响.
- 通过使用四个基准数据集 (德国信用,澳大利亚信用批准,台湾信用卡违约,企业信用风险评估) 验证了优化的SMOTE算法.
主要成果:
- 优化的SMOTE算法在信用风险预测方面显著优于六个比较模型.
- 实现了超过21% (高达38%) 的精度改进,精度超过28% (高达35%),回忆超过31% (高达42%),F1得分约为33% (高达39%).
- 证明了适应性边界调整和优化重量分配在生成更具代表性的合成样本方面的有效性.
结论:
- 优化的SMOTE算法为多元化企业环境中的信用风险预测提供了卓越的解决方案.
- 预测准确度的提高加强了金融机构的风险管理能力和决策能力.
- 该研究通过提高信用风险评估准确度,为金融市场的稳定性做出了贡献.
相关概念视频
Predicting Products: Substitution vs. Elimination
12.3K
When a nucleophile and an alkyl halide react, nucleophilic substitution and β-elimination reactions compete to generate products.
The following factors can influence the mechanisms competing against each other:
The following factors can influence the mechanisms competing against each other:
12.3K
Actuarial Approach
140
The actuarial approach, a statistical method originally developed for life insurance risk assessment, is widely used to calculate survival rates in clinical and population studies. This method accounts for participants lost to follow-up or those who die from causes unrelated to the study, ensuring a more accurate representation of survival probabilities.
Consider the example of a high-risk surgical procedure with significant early-stage mortality. A two-year clinical study is conducted,...
Consider the example of a high-risk surgical procedure with significant early-stage mortality. A two-year clinical study is conducted,...
140
Regression Analysis
6.1K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
6.1K
Aggregates Classification
389
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
389
Multiple Regression
3.2K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.2K
Correlation and Regression
2.0K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
2.0K


