Learning by natural gradient on noncompact matrix-type pseudo-Riemannian manifolds

Simone Fiori1

  • 1Dipartimento di Ingegneria Biomedica, Elettronica e Telecomunicazioni, Facoltà di Ingegneria, Università Politecnica delle Marche, Ancona, Italy. s.fiori@univpm.it

Summary

This study explores natural-gradient optimization on noncompact manifolds. Pseudo-Riemannian metrics offer tractable calculations for learning algorithms, overcoming challenges with traditional Riemannian geometry.

Related Concept Videos

Gradient Vectors and Their Applications01:19

Gradient Vectors and Their Applications

Every point on a topographical map corresponds to a particular elevation, so the landscape can be modeled as a surface whose height depends on horizontal position. From any given location, a hiker may face infinitely many directions, but only one direction produces the fastest possible increase in elevation. This unique route is called the direction of steepest ascent, and in multivariable calculus, it is represented by the gradient vector of the elevation function.The gradient vector points...
Significance of the Gradient Vector01:27

Significance of the Gradient Vector

A surface defined by a function of two variables can be understood by examining how it changes along specific directions. When one variable is held constant, the surface reduces to a curve that reflects variation in the other variable. For example, fixing one variable and moving parallel to a coordinate axis produces a cross-sectional curve. The slope of this curve at a given point represents how the function changes in that particular direction, providing a measure of local steepness.By...
Gradient Fields01:27

Gradient Fields

A gradient field is a vector field derived from a scalar field. A scalar field assigns a single numerical value to every point in space, such as temperature, pressure, or electric potential. The gradient field describes how that value changes from point to point. It gives both the direction of the fastest increase and the rate of change in that direction.For a scalar field f(x, y), the gradient is written as\begin{equation*}\nabla f=\left\langle \jfrac{\partial f}{\partial x},\jfrac{\partial...
Multivariable Functions and Higher Derivatives01:30

Multivariable Functions and Higher Derivatives

A multivariable function assigns a single output value to each ordered set of independent inputs, thereby defining a surface in three-dimensional space. For a function f(x, y), each point (x, y) corresponds to a height z = f(x, y). This geometric interpretation allows systematic analysis of how the output varies as multiple variables change simultaneously. Such functions frequently arise in physical models and optimization problems, where system behavior depends on several interacting...
Routh-Hurwitz Criterion I01:15

Routh-Hurwitz Criterion I

Consider an electrical power grid, where stability is essential to prevent blackouts. The Routh-Hurwitz criterion is a valuable tool for assessing system stability under varying load conditions or faults. By analyzing the closed-loop transfer function, the Routh-Hurwitz criterion helps determine whether the system remains stable.
To apply the Routh-Hurwitz criterion, a Routh table is constructed. The table's rows are labeled with powers of the complex frequency variable s, starting from the...
Introduction to Nonlinear Inequalities01:25

Introduction to Nonlinear Inequalities

Linear and nonlinear inequalities are fundamental for analyzing variable relationships and identifying ranges satisfying specific conditions. A linear inequality involves variables raised only to the first power, resulting in a straight-line graph. This line partitions the coordinate plane into two distinct regions: one that satisfies the inequality and one that does not. Each region represents a set of solutions where the linear relationship holds true under the specified constraint.Nonlinear...