使用XGBoost预测中国股票市场,以最佳权重进行多目标优化.
1School of International Trade and Economics, University of International Business and Economics, Beijing, China.
PeerJ. Computer science
|March 14, 2024
概括
一个新的机器学习模型,最佳权重极端梯度提升 (OW-XGBoost),平衡投资组合风险和回报. 这种人工智能方法增强了金融定量研究和投资策略.
科学领域:
- 人工智能的人工智能
- 机器学习 机器学习
- 金融定量研究 金融定量研究
背景情况:
- 人工智能 (AI) 是一个不断增长的研究领域,在金融领域具有巨大的潜力.
- 机器学习模型为经济和金融分析提供了有价值的工具.
研究的目的:
- 提出和评估最佳重量极端梯度提升 (OW-XGBoost) 模型.
- 为了平衡投资组合的回报和风险,使用一种新的多目标优化方法.
主要方法:
- 开发了OW-XGBoost模型,将标签与最佳重量合并为多目标优化.
- 将模型应用于中国A股数据 (2022年10月 - 2023年4月).
- 在各种市场条件和股票选择中进行了稳定性测试.
主要成果:
- 与基线XGBoost模型 (YL-XGBoost,MLC-XGBoost) 相比,OW-XGBoost在风险控制和回报生成方面表现优越.
- 该模型的整体性能比现有方法更好.
- 稳定性测试证实了不同市场条件,库存池和培训时间的一致性.
结论:
- OW-XGBoost模型为平衡投资组合风险和回报提供了一种有效的方法.
- 该研究为在金融定量研究中整合人工智能和机器学习提供了新的途径.
- 该模型显示,在具有高价值股票和每月培训数据的中度波动的市场中,表现最佳.
相关概念视频
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Regression Analysis
5.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.7K
Goodness-of-Fit Test
3.3K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
3.3K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
End Point Prediction: Gran Plot
322
A Gran plot is used to predict the equivalence volume or endpoint of a potentiometric or acid-base titration without reaching the endpoint. Typically, titration data is collected as a function of the titrant's volume up to a point less than the equivalence volume and then transformed into a linear format. The straight line is extended to the x-axis, indicating the necessary titrant volume to achieve the equivalence point.
For potentiometric titration, the Gran plot is created by plotting...
For potentiometric titration, the Gran plot is created by plotting...
322


