使用改进的NSGA-II-RF算法预测股票价格变动,具有三阶段特征工程过程
Xiaohua Zeng1, Jieping Cai1, Changzhou Liang1
1School of Economics and Trade, Guangzhou Xinhua University, Dongguan, China.
PloS one
|June 28, 2023
概括
本研究介绍了一种改进的多目标优化算法与随机森林 (I-NSGA-II-RF) 股票价格预测. 新方法通过优化功能选择和模型参数来提高准确性并减少计算时间.
科学领域:
- 人工智能的人工智能
- 计算智能是一种计算智能.
- 金融预测 金融预测
背景情况:
- 由于复杂,高维度和非静止的财务数据,准确的股价预测仍然具有挑战性.
- 以前的方法往往忽视了特征工程在提高预测准确性的关键作用.
- 现有的计算智能方法,包括机器学习和深度学习,在处理这些复杂性方面面临局限性.
研究的目的:
- 提出一个改进的多目标优化算法,整合随机森林 (I-NSGA-II-RF) 以提高股票价格预测.
- 通过结合一个新的三阶段特征工程过程来解决以前方法的局限性.
- 减少计算复杂性,提高股票价格预测系统的准确性.
主要方法:
- 开发了一种改进的多目标优化算法 (I-NSGA-II-RF),结合了三阶段的特征工程过程.
- 从两个过的特征选择方法中利用综合信息初始化群体来优化I-NSGA-II算法.
- 采用多染色体混合编码,同时进行特征选择和模型参数优化.
- 将选定的特征和参数输入到随机森林 (RF) 模型中,用于代训练,预测和优化.
主要成果:
- I-NSGA-II-RF算法表现出卓越的性能,实现了最高的平均精度.
- 提出的方法产生了最小的最佳解决方案集,表明了高效的特征选择.
- 实验结果显示,与未经修改的多目标和单目标特征选择算法相比,运行时间显著减少.
- 在准确性和计算速度方面,I-NSGA-II-RF模型的表现优于深度学习模型,同时提供更好的解释性.
结论:
- 通过先进的功能工程和优化,I-NSGA-II-RF算法有效地提高了股票价格预测的准确性和效率.
- 这种方法为深度学习模型提供了有竞争力的替代方案,在高性能的同时提供了可解释性.
- 该研究强调了特征选择在开发强大而准确的财务预测系统中的重要性.
相关概念视频
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Testing a Claim about Standard Deviation
2.5K
A complete procedure to test a claim about population standard deviation or population variance is explained here.
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
2.5K
Response Surface Methodology
190
Response Surface Methodology (RSM) is a collection of statistical and mathematical techniques used to develop, improve, and optimize processes. It is particularly valuable when many input variables or factors potentially influence a response variable.
The process of RSM involves several key steps:
The process of RSM involves several key steps:
190


