Related Experiment Video
Updated: Mar 18, 2026

08:12
A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
Published on: March 1, 2022
3.0K
A Universal Metric of Dataset Similarity for Cross-silo Federated Learning
Ahmed Elhussein1,2, Gamze Gürsoy1,2,3
1Department of Biomedical Informatics, Columbia University, NY, USA.
Summary
Federated Learning (FL) performance suffers with non-IID data. A new privacy-preserving metric predicts beneficial collaborations by analyzing model representations after one training round, improving participant selection.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Data Privacy
Background:
- Federated Learning (FL) facilitates collaborative training without raw data sharing, crucial for privacy-sensitive sectors like healthcare.
- Non-Independent and Identically Distributed (non-IID) datasets significantly degrade FL performance.
- Existing dataset similarity metrics for FL collaboration lack interpretability, require direct data access, or are inefficient.
Purpose of the Study:
- To develop a novel, privacy-preserving metric for assessing client dataset similarity in Federated Learning.
- To predict the impact of client collaboration on model performance in FL settings.
- To provide an actionable tool for selecting optimal participants in cross-silo FL.
Main Methods:
- Proposed a novel metric using model representations from a single federated training round.
- Formulated similarity assessment as an optimal transport problem with a hybrid cost function.
- Ensured privacy using Secure Multiparty Computation (SMC) and Differential Privacy (DP).
Main Results:
- The metric accurately predicts collaboration benefits by analyzing feature-level and label distribution divergence.
- Theoretical analysis links the metric to weight divergence, explaining its predictive power.
- Empirical validation shows strong correlation with weight divergence throughout training.
Conclusions:
- The proposed metric reliably identifies beneficial collaborations in Federated Learning.
- This privacy-preserving approach offers an actionable tool for participant selection in cross-silo FL.
- Early-round model representations effectively predict long-term collaboration outcomes.
Related Concept Videos
Causes of Similarity-Dissimilarity Effect
320
The similarity-dissimilarity effect, a fundamental concept in social psychology, explains how interpersonal similarities and differences influence attraction and social interactions. This effect is supported by three key psychological perspectives: balance theory, social comparison theory, and consensual validation.Balance Theory and Cognitive ConsistencyBalance theory, developed by Fritz Heider, posits that individuals seek cognitive consistency in their relationships. When two people share...
320
One-Way ANOVA: Equal Sample Sizes
4.3K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
4.3K
Improving Translational Accuracy
3.8K
3.8K
Improving Translational Accuracy
15.4K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.4K
Aggregates Classification
1.1K
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
1.1K
Multiple Comparison Tests
4.5K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
4.5K