在回归市场的数据相似性下,隐私意识的数据采集
概括
本研究引入了一个新的数据市场设计,考虑数据相似性和隐私偏好. 它展示了这些因素如何影响分散数据交换中的定价和数据价值.
科学领域:
- 计算机科学 计算机科学
- 经济学 经济学 经济学
- 信息安全 信息安全
背景情况:
- 数据市场使人工智能和机器学习的分散数据交换成为可能.
- 市场设计因各种隐私需求和数据相似性而复杂化.
- 现有的研究往往忽略了数据相似性如何通过信息泄露影响定价和价值.
研究的目的:
- 调查数据相似性和隐私偏好对数据市场设计的影响.
- 为数据采集提出一个新的查询-响应协议.
- 在一个注重隐私的数据市场中建模战略互动.
主要方法:
- 开发了一种使用局部差异隐私 (LDP) 的双方数据采集机制.
- 将市场建模为隐私意识的数据所有者与学习者之间的斯塔克尔伯格游戏.
- 使用数值评估来分析市场参与和数据价值.
主要成果:
- 数据相似性显著影响定价和交易数据的价值.
- 拟议的LDP协议解决了数据交换中的隐私问题.
- 业主和学习者之间的战略互动受到隐私因素的影响.
结论:
- 数据相似性和隐私偏好对于有效的数据市场设计至关重要.
- 该研究为建立更强大,保护隐私的数据市场提供了框架.
- 结果为优化数据价值和参与去中心化市场提供了洞察力.
相关概念视频
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Correlation and Regression
1.2K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
1.2K
Residuals and Least-Squares Property
7.2K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.2K
Quantifying and Rejecting Outliers: The Grubbs Test
1.4K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.4K
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
113
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
113
Regression Analysis
5.5K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.5K


