机器学习模型用于股票市场预测的性能比较与新的投资策略
Azaz Hassan Khan1, Abdullah Shah1, Abbas Ali1
1Department of Electrical Engineering and Computer Science, Jalozai Campus, University of Engineering and Technology, Peshawar, Pakistan.
PloS one
|September 21, 2023
概括
本研究引入了一项新策略,用于增强用于股票市场预测的机器学习 (ML). 拟议的方法显著提高预测准确性,优化回报,同时最大限度地降低财务风险.
科学领域:
- * 计算金融学
- * 金融机器学习
背景情况:
- * 股票市场预测本质上是复杂的,有效市场假设表明完美的预测是不可能的.
- *机器学习 (ML) 为股票市场预测准确性提供了潜在的改进.
- * 传统的ML验证方法可能无法完全捕捉现实金融场景中的业绩.
研究的目的:
- * 提出一种新的策略,以提高金融市场中ML模型的预测效率.
- * 通过传统和拟议的方法来评估九个ML模型的性能.
- * 评估ML模型的风险回报概况,而不仅仅是简单的准确度指标.
主要方法:
- *通过使用历史1天的股票市场数据,训练和验证了9个ML模型.
- *开发了一种新的培训方法,并应用于这些模型.
- * 用金融市场模拟来评估风险,最大提款和回报.
主要成果:
- *在传统方法下,物流回归产生了最高的准确性 (85.51%).
- * 拟议的策略导致Random Forest获得最高准确率 (91.27%),超过了其他模型,如XG Boost,ADA Boost和ANN.
- *模拟结果表明,拟议的策略提高了股票市场的回报率,并降低了相关风险.
结论:
- *新的策略显著提高了ML模型对股票市场方向的预测准确度.
- *仅在分类报告上评估ML模型是不够的;风险回报模拟是至关重要的.
- * 拟议的方法为基于ML的股票市场预测提供了更强大的方法,平衡业绩和风险.
更多相关视频
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
8.3K
07:35Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
7.5K
相关概念视频
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Regression Analysis
5.8K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.8K
Pharmacokinetic Models: Comparison and Selection Criterion
96
Physiological and compartmental models are valuable tools used in studying biological systems. These models rely on differential equations to maintain mass balance within the system, ensuring an accurate representation of the dynamic processes at play.
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
96
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Correlation and Regression
1.3K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
1.3K
Goodness-of-Fit Test
3.4K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
3.4K
