Related Experiment Video
Updated: Nov 4, 2025

Exploring Caspase Mutations and Post-Translational Modification by Molecular Modeling Approaches
Published on: October 13, 2022
CoolMomentum: a method for stochastic optimization by Langevin dynamics with simulated annealing
Oleksandr Borysenko1, Maksym Byshkin2
1National Science Center "Kharkiv Institute of Physics and Technology", Kharkiv, 61108, Ukraine. alessandro.borisenko@gmail.com.
We introduce CoolMomentum, a novel stochastic optimization method inspired by physics. By gradually decreasing momentum, it mimics Simulated Annealing, achieving high accuracy in deep learning tasks like Resnet-20 and Efficientnet-B0.
Area of Science:
- Computational Mathematics
- Machine Learning
- Statistical Physics
Background:
- Deep learning necessitates global optimization of non-convex functions with multiple local minima.
- This challenge mirrors problems in physical simulations, often addressed by Langevin dynamics with Simulated Annealing.
Purpose of the Study:
- To bridge insights from physical simulations to non-convex stochastic optimization in machine learning.
- To propose a novel optimization method, CoolMomentum, inspired by physical principles.
Main Methods:
- Discretizing the Langevin equation to derive a coordinate updating rule.
- Establishing an equivalence between decreasing the momentum coefficient and Simulated Annealing (slow cooling).
- Implementing the proposed CoolMomentum optimization algorithm.
Main Results:
- The coordinate updating rule derived from the Langevin equation is equivalent to the Momentum optimization algorithm.
- A gradual decrease of the momentum coefficient is shown to be analogous to Simulated Annealing.
- CoolMomentum achieves high accuracy on Resnet-20 (Cifar-10) and Efficientnet-B0 (ImageNet).
Conclusions:
- The analogy between physical simulations and machine learning optimization offers valuable insights.
- CoolMomentum presents a novel and effective approach for stochastic optimization in deep learning.
- This method demonstrates strong performance on complex deep learning models and datasets.
More Related Videos
10:36Author Spotlight: Optimization of Airflow Velocities in Battery Cooling Systems for Enhanced Thermal Performance and Reduced Energy Consumption
Published on: November 3, 2023
11:03An Analog Macroscopic Technique for Studying Molecular Hydrodynamic Processes in Dense Gases and Liquids
Published on: December 4, 2017
Related Concept Videos
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Principle of Linear Impulse and Momentum for a Single Particle: Problem Solving
Angular Momentum: Single Particle
Stability of Equilibrium Configuration: Problem Solving
Problem-solving in the context of the stability of equilibrium configuration...
Maxwell-Boltzmann Distribution: Problem Solving
This distribution function f(v) is defined by saying that the expected number N (v1,v2) of particles with speeds between v1 and v2 is given by
Entropy Change in Reversible Processes
The statement can be further generalized to prove that entropy is a state function. Take a cyclic process between any two points on a p-V diagram.