通过无生成器的数据生成进行无数据的知识蒸,用于非IID联合学习的非IID联合学习
Siran Zhao1, Tianchi Liao2, Lele Fu3
1Sun Yat-sen University, School of computer science and engineering, Guangzhou, China.
概括
联合学习 (FL) 面临非IID数据挑战. 我们的FedF^2DG方法使用本地模型进行数据生成和知识蒸,改善了没有代理数据或生成器的全球模型性能.
科学领域:
- 人工智能的人工智能
- 机器学习 机器学习
- 分布式系统 分布式系统
背景情况:
- 非IID数据异质性是联邦学习 (FL) 中的一个主要挑战,导致本地模型漂移和性能下降.
- 现有的FL知识蒸方法通常依赖于代理数据集或数据生成器,这些数据集或数据生成器并不总是可用或可靠的.
- 生成数据和服务器依赖的生成器的不稳定性限制了当前方法的有效性.
研究的目的:
- 为非IID FL 提出一种新的无数据知识蒸方法,命名为 FedF^2DG.
- 在FL场景中克服代理数据集和数据生成器的局限性.
- 提高全球模型在异质联合环境中的性能.
主要方法:
- FedF^2DG使用本地模型为每个客户端生成伪数据集,从而实现无数据知识蒸.
- 引入了一个规范化术语,通过利用本地和全球模型之间的分歧来生成硬样本.
- 数据生成原则根据客户端状态自适应地控制伪数据集标签的分布和数量.
主要成果:
- 在非IID设置中,FedF^2DG显著优于最先进的FL方法.
- 该方法通过自适应的伪数据集生成,有效地提取更多的客户端知识.
- 实验表明,当FedF^2DG作为FedAvg和FedProx等现有FL算法的插件时,其性能得到了改进.
结论:
- FedF^2DG为非IID联合学习提供了一个有效的无数据解决方案.
- 拟议的方法在异质数据分布中提高了全球模型的准确性和稳定性.
- FedF^2DG提供了一种灵活,高性能的方法来改进FL系统.
相关概念视频
Data: Types and Distribution
709
In biostatistics, data are the observations collected for analysis. There are two main types: parametric and non-parametric. Parametric data, which include continuous (e.g., weight) and discrete numerical data (e.g., number of tablets), assume a particular distribution pattern, often the normal distribution. Non-parametric data do not adhere to a specific distribution and typically comprise nominal (e.g., gender) and ordinal categorical data (e.g., pain scale ratings).
Distributions in...
Distributions in...
709
Random Sampling Method
11.0K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. Data are the result of sampling from a population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest. Among the various sampling methods used by...
11.0K
Cluster Sampling Method
11.8K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.8K
Data Collection by Observations
11.9K
Data collection refers to a systematic way of obtaining, observing, measuring, and analyzing accurate information. Observational studies are one of the most widely used methods of data collection. It involves collecting data by observing the behavior and physical characteristics of a sample without making any modifications to the sample.
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
11.9K
Randomized Experiments
6.8K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
6.8K
Data Collection by Experiments
24.0K
Data collection is a systematic method of obtaining, observing, measuring, and analyzing accurate information. An experimental study is a standard method of data collection that involves the manipulation of the samples by applying some form of treatment prior to data collection. It refers to manipulating one variable to determine its changes on another variable. The sample subjected to treatment is known as “experimental units.”
An example of the experimental method is a public...
An example of the experimental method is a public...
24.0K


