Related Experiment Video
Updated: Aug 22, 2025

Author Spotlight: Efficient Image Recognition Using Directional Gradient Histogram Technique and Support Vector Machines
Published on: January 5, 2024
Convergence Analysis of Distributed Gradient Descent Algorithms With One and Two Momentum Terms
Abstract:
For the centralized optimization, it is well known that adding one momentum term (also called the heavy-ball method) can obtain a faster convergence rate than the gradient method. However, for the distributed counterpart, there is quite few results about the effect of added momentum terms on the convergence rate. This article is aimed at studying the issue in the distributed setup, where N agents minimize the sum of their individual cost functions using local communication over a network. The cost functions are twice continuously differentiable. We first study the algorithm with one momentum term and develop a distributed heavy-ball (D-HB) method by adding one momentum term on to the distributed gradient algorithm. By borrowing tools from the control theory, we provide a simple convergence proof and an explicit expression of the optimal convergence rate. Furthermore, we consider adding two momentum terms case and propose a distributed double-heavy-ball (D-DHB) method. We show that adding one momentum term allows faster convergence while adding two momentum terms does not perform any superiorities. Finally, simulation examples are given to illustrate our findings.
Related Concept Videos
Collisions in Multiple Dimensions: Introduction
Collisions in Multiple Dimensions: Problem Solving
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
Gradient and Del Operator
Conservation of Momentum: Introduction
Principle of Linear Impulse and Momentum for a Single Particle
Delving...
Linear Momentum in Control Volume

