数据集相似性的通用指标,用于跨 silo 联合学习
Ahmed Elhussein1,2, Gamze Gürsoy1,2,3
1Department of Biomedical Informatics, Columbia University, NY, USA.
概括
联合学习 (FL) 的表现受到非IID数据的影响. 一个新的隐私保护指标通过在一轮培训后分析模型表示来预测有益的合作,从而改善了参与者选择.
科学领域:
- 人工智能的人工智能
- 机器学习 机器学习
- 数据 隐私 数据 隐私 数据
背景情况:
- 联合学习 (FL) 促进了无需共享原始数据的协作培训,这对于医疗保健等对隐私敏感的领域至关重要.
- 非独立和相同分布 (非IID) 数据集显著降低了FL的性能.
- 现有的数据集相似度指标对于FL协作缺乏解释性,需要直接访问数据,或者是低效的.
研究的目的:
- 开发一种新的,保护隐私的指标,用于评估客户数据集在联合学习中的相似性.
- 预测客户端协作对FL设置中的模型性能的影响.
- 为选择最佳参与者提供可操作的工具.
主要方法:
- 提出了一种使用单个联合培训轮模型表示的新型度量.
- 制定了相似性评估,作为一种混合成本函数的最佳运输问题.
- 使用安全多方计算 (SMC) 和差分隐私 (DP) 确保隐私.
主要成果:
- 该指标通过分析功能级别和标签分发差异来准确预测协作效益.
- 理论分析将度量与重量分歧联系起来,解释了它的预测能力.
- 经验验证表明,在整个训练过程中,体重差异与体重差异有很强的相关性.
结论:
- 拟议的指标可靠地识别了联邦学习中的有益合作.
- 这种维护隐私的方法提供了一个可操作的工具,用于参与者选择跨 Silo FL.
- 早期的模型表示有效地预测了长期合作成果.
相关概念视频
Causes of Similarity-Dissimilarity Effect
320
The similarity-dissimilarity effect, a fundamental concept in social psychology, explains how interpersonal similarities and differences influence attraction and social interactions. This effect is supported by three key psychological perspectives: balance theory, social comparison theory, and consensual validation.Balance Theory and Cognitive ConsistencyBalance theory, developed by Fritz Heider, posits that individuals seek cognitive consistency in their relationships. When two people share...
320
One-Way ANOVA: Equal Sample Sizes
4.3K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
4.3K
Improving Translational Accuracy
3.8K
3.8K
Improving Translational Accuracy
15.4K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.4K
Aggregates Classification
1.1K
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
1.1K
Multiple Comparison Tests
4.5K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
4.5K
