在流媒体和滑动窗模型中最大限度地最大化公平的最大限度的多样性
Yanhao Wang1, Francesco Fabbri2, Michael Mathioudakis3
1School of Data Science and Engineering, East China Normal University, Shanghai 200062, China.
Entropy (Basel, Switzerland)
|July 29, 2023
概括
我们开发了快速算法,以最大限度地提高数据流中的公平多样性. 我们的方法有效地选择多样化的子集,同时确保群体的代表性,优于以前的方法.
科学领域:
- 计算机科学 计算机科学
- 数据科学数据科学数据科学
- 算法设计 算法设计
背景情况:
- 多样性最大化对于推系统等应用至关重要.
- 在数据分析中,公平性约束越来越重要.
- 现有的公平多样性算法对于数据流是低效的.
研究的目的:
- 开发高效的算法,以实现公平的最大-最小多样性最大化.
- 为了应对流媒体和滑动窗数据模型的挑战.
- 通过纳入集团代表制约来确保公平.
主要方法:
- 为仅插入流媒体模型设计了近似算法.
- 开发了滑窗模型的近似算法.
- 在真实世界和合成数据集上评估性能.
主要成果:
- 实现了与离线算法可比的解决方案质量.
- 显示了显著的速度改进 (数量级更快).
- 算法对流媒体和滑动窗口设置都非常有效.
结论:
- 提出的算法有效地平衡了数据流中的多样性和公平性.
- 这些方法为实时公平多样性最大化提供了实际解决方案.
- 这项工作推进了公平的数据总结和选择的最新技术.
相关概念视频
Friedman Two-way Analysis of Variance by Ranks
240
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
240
Expected Frequencies in Goodness-of-Fit Tests
2.6K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.6K
Linear Approximation in Frequency Domain
111
Linear systems are characterized by two main properties: superposition and homogeneity. Superposition allows the response to multiple inputs to be the sum of the responses to each individual input. Homogeneity ensures that scaling an input by a scalar results in the response being scaled by the same scalar.
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
111
Frequency-dependent Selection
22.1K
When the fitness of a trait is influenced by how common it is (i.e., its frequency) relative to different traits within a population, this is referred to as frequency-dependent selection. Frequency-dependent selection may occur between species or within a single species. This type of selection can either be positive—with more common phenotypes having higher fitness—or negative, with rarer phenotypes conferring increased fitness.
22.1K
Central Limit Theorem
15.3K
The central limit theorem, abbreviated as clt, is one of the most powerful and useful ideas in all of statistics. The central limit theorem for sample means says that if you repeatedly draw samples of a given size and calculate their means, and create a histogram of those means, then the resulting histogram will tend to have an approximate normal bell shape. In other words, as sample sizes increase, the distribution of means follows the normal distribution more closely.
The sample size, n, that...
The sample size, n, that...
15.3K
Sampling Distribution
13.0K
Given simple random samples of size n from a given population with a measured characteristic such as mean, proportion, or standard deviation for each sample, the probability distribution of all the measured characteristics is called a sampling distribution. How much the statistic varies from one sample to another is known as the sampling variability of a statistic. You typically measure the sampling variability of a statistic by its standard error. The standard error of the mean is an example...
13.0K


