STTM:一种有效的方法来估计新闻对股票运动方向的影响
Aleksei Riabykh1, Denis Surzhko1, Maxim Konovalikhin1
1Department of Data Analysis and Modeling, VTB Bank, Moscow, Russia.
PeerJ. Computer science
|June 22, 2023
概括
这项研究引入了一个新的主题建模框架来分析金融新闻,成功预测股票市场的回报. 该方法通过揭示经济信息和股票价格之间的重要关系来增强交易策略.
科学领域:
- 计算金融是指计算金融.
- 自然语言处理自然语言处理.
- 计量经济学 计量经济学 计量经济学
背景情况:
- 金融新闻显著影响股票市场的行为,但将文本数据与股票价格变动联系起来的算法发展不足.
- 主题建模是一种缩小维度的技术,对于分析财务文本和预测市场趋势,它仍然在很大程度上未被探索.
研究的目的:
- 开发和评估一个主题建模框架,以评估金融新闻和股价之间的关系,以提高交易收益.
- 证明该框架能够检测回报预测信号,并优于现有模型的性能.
主要方法:
- 使用了来自俄罗斯三个媒体 (商报,维多莫斯蒂,RIA Novosti) 的197,678篇经济文章的数据集.
- 应用主题建模来预测2013年至2021年的39个高度流动的俄罗斯股票的时间序列.
- 使用夏普比率,年回报率和格兰杰因果关系测试来评估模型性能.
主要成果:
- 主题建模框架检测到重要的回报预测信号,在夏普比率和年回报方面表现优于26个现有模型.
- 在超过70%的投资组合股票中发现了显著的格兰杰因果关系.
- 该方法产生了可解释的结果,不需要特定领域的字典,并且可以对单个时间序列进行校准.
结论:
- 提出的主题建模框架有效地将金融新闻与股票市场行为联系起来,为交易策略和分析提供实际好处.
- 该方法的可解读性,适应性和证明的成功表明其潜在的可转移到其他市场,包括欧洲股票市场.
相关概念视频
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Testing a Claim about Standard Deviation
2.5K
A complete procedure to test a claim about population standard deviation or population variance is explained here.
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
2.5K
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Residuals and Least-Squares Property
7.4K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.4K
Central Tendency: Analysis
175
Measures of central tendency are tools used in biostatistics to identify the average or center of a dataset. They offer a single representative value for understanding and summarizing data distribution.
The mean is one such measure, calculated by totaling all values in a dataset and dividing by the number of values. For instance, the mean blood pressure reading (120, 130, 140, 150) would be 135. However, the mean can be affected by extreme values or outliers.
The median, another measure,...
The mean is one such measure, calculated by totaling all values in a dataset and dividing by the number of values. For instance, the mean blood pressure reading (120, 130, 140, 150) would be 135. However, the mean can be affected by extreme values or outliers.
The median, another measure,...
175
The Anchoring-and-Adjustment Heuristic
7.3K
In order to make good decisions, we use our knowledge and our reasoning. Often, this knowledge and reasoning is sound and solid. However, sometimes, we are swayed by biases or by others manipulating a situation. For example, let’s say you and three friends wanted to rent a house and had a combined target budget of $1,600. The realtor shows you only very run-down houses for $1,600 and then shows you a very nice house for $2,000. Might you ask each person to pay more in rent to get the...
7.3K


