通过基于预测的沙普利值来提高对商品的自动编码因子的解释性.
Roy Cerqueti1,2, Antonio Iovanella3, Raffaele Mattera4
1Department of Social and Economic Sciences, Sapienza University of Rome, P.le Aldo Moro 5, 00185, Rome, Italy. roy.cerqueti@uniroma1.it.
Scientific reports
|August 23, 2024
概括
这项研究提高了使用Shapley值的金融自动编码器可解释性. 它基于预测准确性来确定大宗商品市场的关键非线性潜在因素.
科学领域:
- 机器学习 机器学习
- 计量经济学 计量经济学
- 金融建模金融建模
背景情况:
- 自动编码器是机器学习中的强大的维度减小工具,类似于主要组件分析 (PCA).
- 由于灵活性和性能,它们在非线性因子模型的融资中的应用正在增长.
- 一个主要的限制是与PCA相比,可解释性降低.
研究的目的:
- 提高非线性因子模型中自动编码器的可解释性.
- 引入一种新的Shapley基于价值的方法来评估潜在因素的相关性.
- 确定金融市场,特别是商品市场中重要的非线性潜伏因素.
主要方法:
- 使用Shapley值来量化每个非线性潜伏因子的贡献.
- 实施基于预测的沙普利值方法,用于样本之外的准确度测量.
- 将该方法应用于商品市场的因子增大模型.
主要成果:
- 沙普利值方法有效地衡量非线性潜伏因子的相关性.
- 根据预测业绩,确定了对个别商品最有影响力的潜在因素.
- 证明了基于自动编码器的因子模型的可解释性得到改善.
结论:
- 沙普利值提供了一个强大的方法来提高金融因素建模中自动编码器的可解释性.
- 基于预测的方法为非样本预测提供了对潜在因素重要性有价值的见解.
- 这种技术对于理解商品等市场的复杂动态尤其有用.
相关概念视频
Variation
6.8K
An important characteristic of any set of data is the variation in the data. In some data sets, the data values are concentrated closely near the mean; in other data sets, the data values are more widely spread out from the mean. The most common measure of variation, or spread, is the standard deviation, which is the square root of variance.
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
6.8K
Extraction: Partition and Distribution Coefficients
2.3K
The distribution law or Nernst's distribution law is the law that governs the distribution of a solute between two immiscible solvents. This law, also known as the partition law, states that if a solute is added to the mixture of two immiscible solvents at a constant temperature, the solute is distributed between the two solvents in such a way that the ratio of solute concentrations in the solvents remains constant at equilibrium.
For extracting a solute from an aqueous phase into an...
For extracting a solute from an aqueous phase into an...
2.3K
Residual Plots
4.6K
A residual plot is a statistical representation of data used to analyze correlation and regression results. It helps verify the requirements for drawing specific conclusions about correlation and regression. To obtain the residual plot, first, the residual for each data value is calculated, which is simply the vertical distance between the observed and the predicted value obtained from the regression equation.
When the residual values are plotted against the variable x, it is called a residual...
When the residual values are plotted against the variable x, it is called a residual...
4.6K
Variability: Analysis
135
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
The range is a simple measure of variability, indicating the difference between the highest and...
135
Random Error
848
Random or indeterminate errors originate from various uncontrollable variables, such as variations in environmental conditions, instrument imperfections, or the inherent variability of the phenomena being measured. Usually, these errors cannot be predicted, estimated, or characterized because their direction and magnitude often vary in magnitude and direction even during consecutive measurements. As a result, they are difficult to eliminate. However, the aggregate effect of these errors can be...
848
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K


