Related Experiment Videos
MoFedAGR: Mitigating client drift with adaptive gradient regularization and global momentum in federated learning
Xiang Wang1, Lei Tian1, Jiahao Gan1
1Zhejiang Key Laboratory of Intelligent Education Technology and Application, Zhejiang Normal University, Jinhua, 321004, China; School of Computer Science and Technology, Zhejiang Normal University, Jinhua, 321004, China.
Abstract:
Federated learning is a novel distributed machine learning framework with privacy-protection, yet it is vulnerable to the effects of heterogeneous data. Heterogeneous data drive client models that overfit local datasets and depart from the global optimum during local training, which is termed client drift. To address the impact of client drift, we approach this issue from the perspectives of optimization and generalization. We comprehensively considering the effects of client drift during the training process, and quantifying it as the aggregation error. We first propose adaptive gradient regularization, which is based on gradient regularization and further and applies different regularization strengths to each parameter based on the magnitude of the parameter variance between the local model and the global model, thereby mitigating the performance degradation caused by aggregation error and helping model converge to a flatter minimum. In order to obtain the variance between local and global models to compute our adaptive gradient regularization term, we introduce global momentum from the server side as the approximation of global gradient and further utilize it as a gradient correction term. Next, we propose MoFedAGR, which combines gradient correction term and adaptive gradient regularization term, helping client models converge to a consistent flat minimum. We have provided the theoretical convergence bounds of the algorithm we proposed. Furthermore, experiments on several image classification datasets demonstrate that our algorithm significantly improves model performance while exhibiting strong generalization capabilities.
Related Concept Videos
Maximizing the Directional Derivative
Gradient Vectors and Their Applications
Regression Toward the Mean
Differential Leveling
Observational Learning
Reducing Line Loss
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss in...