MFedBN:用梯度基聚和先进分布偏建模解决数据异质性问题
Kinda Mreish1,2, Evgenia Novikova1, Mikhail Chaplygin1
1Faculty of Computer Science and Technology, Saint Petersburg Electrotechnical University "LETI", Saint Petersburg 197376, Russia.
Sensors (Basel, Switzerland)
|December 11, 2025
概括
联合学习 (FL) 绩效与非IID数据下降. 一种新的方法,MFedBN,在这些具有挑战性的环境中改进了聚合策略和模型稳定性,增强了保护隐私的应用程序.
科学领域:
- 人工智能的人工智能
- 机器学习 机器学习
- 数据科学数据科学数据科学
背景情况:
- 联合学习 (FL) 促进边缘设备的协作培训,保护数据隐私.
- 对于非独立且相同分布 (非IID) 数据,FL的性能显著下降,阻碍了现实世界的应用.
- 现有的聚合策略在各种非IID数据条件下难以保持性能.
研究的目的:
- 在非IID FL环境中评估聚合策略.
- 提出创新方法,用于生成代表各种非IID类型的偏斜数据集.
- 引入和验证一个新的聚合策略,MFedBN,以提高FL的性能.
主要方法:
- 开发了特征分布倾斜,标签分布倾斜,相同标签/不同特征倾斜以及相同特征/不同标签倾斜的数据生成技术.
- 通过本地批量规范化 (MFedBN) 引入修改后联合,采用服务器端渐变式更新,具有不同的学习速度.
- 对商用车辆传感器和NF-UNSW-NB15数据集进行了实验.
主要成果:
- 在大多数场景中,MFedBN的表现超过了基线FedBN.
- 实现了高测试准确度:在商用车辆传感器上达到85%,在NF-UNSW-NB15.5上达到99.98%.
- 在异质的FL环境中证明了MFedBN的更好的趋同稳定性和概括性.
结论:
- 在各种非IID数据偏差中,MFedBN有效地提高了FL的性能和稳定性.
- 拟议的数据生成方法为评估非IID环境中的FL策略提供了一个平台.
- 这项工作促进了隐私保护FL在现实世界物联网监控和网络入侵检测中的应用.
相关概念视频
Skewness
17.6K
The measures of central tendency calculated from a data set may not reveal much about its intrinsic distribution. If a plot is made of the data set’s values, the mean and the median may not only differ, but also the plot may have more values on one side of the central tendencies. Such a data set is said to be skewed towards that side.
The longer the tail of the plot on one side, the more skewed it is. The skewness of a data set’s values suggests that the measures of central tendency...
The longer the tail of the plot on one side, the more skewed it is. The skewness of a data set’s values suggests that the measures of central tendency...
17.6K
Types of Skewness
17.5K
If the frequency distribution of a data set is more inclined towards smaller or larger values, the distribution is said to be skewed. If data values are skewed to the right, then the distribution is called positively skewed. Conversely, if the plot is skewed to the left, the distribution is called negatively skewed.
For instance, in the middle of a pandemic, the geographical distribution of vaccine coverage may be positively skewed towards populations in the global north countries. However,...
For instance, in the middle of a pandemic, the geographical distribution of vaccine coverage may be positively skewed towards populations in the global north countries. However,...
17.5K
Data: Types and Distribution
1.5K
In biostatistics, data are the observations collected for analysis. There are two main types: parametric and non-parametric. Parametric data, which include continuous (e.g., weight) and discrete numerical data (e.g., number of tablets), assume a particular distribution pattern, often the normal distribution. Non-parametric data do not adhere to a specific distribution and typically comprise nominal (e.g., gender) and ordinal categorical data (e.g., pain scale ratings).
Distributions in...
Distributions in...
1.5K
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
223
Pharmacokinetic models are mathematical constructs that represent and predict the time course of drug concentrations in the body, providing meaningful pharmacokinetic parameters. These models are categorized into compartment, physiological, and distributed parameter models.
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
223
Distributions to Estimate Population Parameter
5.0K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
5.0K
Choosing Between z and t Distribution
3.5K
The z and the Student t distribution estimate the population mean using the sample mean and standard deviation. However, to decide which distribution to use for a calculation, one needs to determine the sample size, the nature of the distribution, and whether the population standard deviation is known. If the population standard deviation is known and the population is normally distributed, or if the sample size is greater than 30, the z distribution is preferred. The Student t distribution is...
3.5K


