Related Experiment Video
Updated: Jun 10, 2025

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
Published on: March 1, 2022
FedDSS: A data-similarity approach for client selection in horizontal federated learning
Tuong Minh Nguyen1, Kim Leng Poh1, Shu-Ling Chong2
1Department of Industrial Systems Engineering and Management, National University of Singapore, Singapore, 117576, Singapore.
Federated Data Similarity Selection (FedDSS) improves federated learning by selecting clients based on data similarity, enhancing model performance and convergence speed for better sepsis prediction.
Area of Science:
- Machine Learning
- Distributed Systems
- Healthcare Informatics
Background:
- Federated learning (FL) enables collaborative model training without sharing sensitive data, addressing healthcare data challenges.
- Non-independent and identically distributed (non-i.i.d.) data across clients causes model divergence and performance degradation in FL.
- Existing FL methods struggle with heterogeneous data distributions, impacting model accuracy.
Purpose of the Study:
- To introduce FedDSS (Federated Data Similarity Selection), a novel FL framework designed to mitigate performance issues caused by non-i.i.d. data.
- To enhance model convergence and predictive accuracy in FL settings by employing a data-similarity client selection strategy.
- To ensure client data privacy is maintained throughout the federated learning process.
Main Methods:
- FedDSS utilizes a statistical data similarity metric, an N-similar-neighbor network, and a network-based client selection strategy.
- Performance was evaluated against FedAvg using pediatric sepsis datasets (PICD, MIMICIII) in both i.i.d. and non-i.i.d. scenarios.
- Key metrics included average loss, true positive rate (TPR), and selection fairness (entropy).
Main Results:
- FedDSS demonstrated faster convergence and higher True Positive Rates (TPR) compared to FedAvg in both i.i.d. and non-i.i.d. settings.
- On PICD, FedDSS achieved higher TPR earlier than FedAvg. On MIMICIII, FedDSS showed significant loss reduction and faster TPR achievement.
- FedDSS exhibited reasonable selection fairness across different data distributions.
Conclusions:
- FedDSS effectively improves federated learning performance by enabling faster convergence and achieving desired TPR with fewer communication rounds.
- The proposed data-similarity approach enhances sepsis prediction accuracy within the FL framework.
- FedDSS offers a privacy-preserving solution to the challenges posed by non-i.i.d. data in healthcare FL applications.
Related Concept Videos
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
One-Way ANOVA: Unequal Sample Sizes
Stratified Sampling Method
To choose a stratified sample, divide the population into groups called strata and then take a...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Relationship Formation
Modeling and Similitude

