一个过度调整的NLP股票回报预测的概率指标.
1Department of Applied Mathematics and Statistics, Johns Hopkins University, Baltimore, MD, United States.
Frontiers in artificial intelligence
|March 6, 2025
概括
本研究引入了用于自然语言处理 (NLP) 库存预测的作调整的概率测量. 这种新的方法通过纠正新闻情绪偏见和转变来提高预测准确性.
科学领域:
- 量化金融 量化金融
- 自然语言处理 (NLP) 是一种自然语言处理.
- 计算金融是指计算金融.
背景情况:
- 传统的财务预测模型经常与新闻情绪的动态性质作斗争.
- 现有的自然语言处理 (NLP) 技术可能无法充分捕捉市场情绪的细微差别.
- 资产定价的融资提供了诸如变化概率度量的工具,这些工具尚未完全融入NLP预测中.
研究的目的:
- 引入一种新的作调整的概率指标,以改善股票回报率和波动性预测.
- 开发一种新的情绪评分方程,以计算一天内新闻影响.
- 将概率测量概念的应用从资产定价扩展到NLP驱动的财务预测.
主要方法:
- 开发了一个新的情绪评分方程来量化一天内新闻的影响.
- 作调整的概率测量是通过重新分配概率权重来构建的.
- 该方法解决了新闻偏见,记忆,重量和情绪方向的转变.
- 预测适用于美国半导体股票.
主要成果:
- 拟议的作调整的概率指标提高了股票回报率和波动性预测的准确性.
- 新闻情绪得分有效地捕捉了当天新闻的影响.
- 该方法证明了金融概率测量的成功扩展到NLP预测中.
结论:
- 过度调整的概率测量为基于NLP的财务预测提供了重大进展.
- 这种方法为将新闻情绪纳入预测模型提供了更强大的方法.
- 该研究强调了将先进的金融数学工具与NLP集成为市场分析的潜力.
相关概念视频
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K
Confidence Coefficient
7.5K
The confidence coefficient is also known as the confidence level or degree of confidence. It is the percent expression for the probability, 1-α, that the confidence interval contains the true population parameter assuming that the confidence interval is obtained after sufficient unbiased sampling; for example, if the CL = 90%, then in 90 out of 100 samples the interval estimate will enclose the true population parameter. Here α is the area under the curve, distributed equally under...
7.5K
Statistical Hypothesis Testing
1.8K
Hypothesis testing is a critical statistical procedure facilitating informed, evidence-based decisions. It begins with a hypothesis, which is a tentative explanation, or a prediction about a population parameter. This hypothesis can be either a null hypothesis (H0), indicating no effect or difference, or an alternative hypothesis (Ha), suggesting an effect or difference.
Statistical significance measures the probability that an observed result occurred by chance. If this probability, known as...
Statistical significance measures the probability that an observed result occurred by chance. If this probability, known as...
1.8K
Probability Histograms
11.0K
A probability histogram is a visual representation of a probability distribution. Similar a typical histogram, the probability histogram consists of contiguous (adjoining) boxes. It has both a horizontal axis and a vertical axis. The horizontal axis is labeled with what the data represents. The vertical axis is labeled with probability. Each rectangular bar in the histogram is 1 unit wide, which suggests that the area under each bar equals the probability, P(x), where x is 1, 2, 3, and so on.
11.0K
Confidence Intervals
6.1K
An unbiased point estimate is often insufficient to predict a population estimate, such as population mean or population proportion. In this scenario, a confidence interval is used. A confidence interval is an estimate similar to a sample proportion. However, unlike the point estimate which is a single value, the confidence interval contains a range of values. These values have lower and upper limits, known as confidence limits, and can be designated as L1 and L2, respectively.
A...
A...
6.1K
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K


