Tuning-Free Coreset Markov Chain Monte Carlo via Hot DoG
Naitong Chen1, Jonathan H Huggins2, Trevor Campbell1
1Department of Statistics, University of British Columbia, Vancouver, BC, Canada.
Summary
We introduce Hot-start Distance over Gradient (Hot DoG), a novel method for training Bayesian coreset weights. Hot DoG eliminates the need for learning rate tuning in Coreset Markov chain Monte Carlo (MCMC), improving posterior approximation quality.
Area of Science:
- Computational Statistics
- Machine Learning
- Bayesian Inference
Background:
- Bayesian coresets offer computational savings by approximating large datasets with smaller, weighted subsets.
- Current Coreset Markov chain Monte Carlo (MCMC) methods rely on stochastic gradient optimization, which is sensitive to learning rate hyperparameters.
- Suboptimal learning rates can degrade the quality of the coreset and its resulting posterior approximation.
Purpose of the Study:
- To develop a learning-rate-free optimization procedure for training Bayesian coreset weights.
- To enhance the robustness and reduce user tuning effort in Coreset MCMC algorithms.
- To improve the quality of posterior approximations obtained using Bayesian coresets.
Main Methods:
- Propose Hot-start Distance over Gradient (Hot DoG), a novel learning-rate-free stochastic gradient optimization algorithm.
- Provide a theoretical analysis of the convergence properties of coreset weights trained with Hot DoG.
- Empirically evaluate Hot DoG against existing learning-rate-free methods and adaptive optimizers like ADAM.
Main Results:
- Hot DoG successfully trains coreset weights without requiring manual learning rate selection.
- Theoretical analysis confirms the convergence of coreset weights generated by Hot DoG.
- Empirical results show Hot DoG yields superior posterior approximations compared to other learning-rate-free methods and is competitive with tuned ADAM.
Conclusions:
- Hot DoG offers a robust and user-friendly alternative for training Bayesian coresets within Coreset MCMC.
- The proposed method achieves high-quality posterior approximations without hyperparameter tuning.
- This advancement reduces computational cost and complexity in Bayesian inference with large datasets.
Related Concept Videos
Maxam-Gilbert Sequencing
12.5K
In the same year as the discovery of the Sanger sequencing method, another group of scientists, Allan Maxam and Walter Gilbert, demonstrated their chemical-cleavage method for DNA sequencing. The Maxam-Gilbert method relies on using different chemicals that can cleave the DNA sequence at specific sites, the separation of resulting DNA fragments of variable size using electrophoresis, and deciphering the DNA sequence from the resulting gel bands.
Challenges of the Maxam-Gilbert Method
The...
Challenges of the Maxam-Gilbert Method
The...
12.5K
Random Variables
17.2K
A random variable is a single numerical value that indicates the outcome of a procedure. The concept of random variables is fundamental to the probability theory and was introduced by a Russian mathematician, Pafnuty Chebyshev, in the mid-nineteenth century.
Uppercase letters such as X or Y denote a random variable. Lowercase letters like x or y denote the value of a random variable. If X is a random variable, then X is written in words, and x is given as a number.
For example, let X = the...
Uppercase letters such as X or Y denote a random variable. Lowercase letters like x or y denote the value of a random variable. If X is a random variable, then X is written in words, and x is given as a number.
For example, let X = the...
17.2K
Randomized Experiments
8.8K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
8.8K
Per-Unit Sequence Models
404
An ideal Y-Y transformer, grounded through neutral impedances, displays per-unit sequence networks akin to those of a single-phase ideal transformer when subjected to balanced positive- or negative-sequence currents. These currents do not produce neutral currents, and their associated voltage drops.
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
404
Sequence Networks of Rotating Machines
467
A Y-connected synchronous generator, grounded through a neutral impedance, is designed to produce balanced internal phase voltages with only positive-sequence components. The generator's sequence networks include a source voltage that is exclusively in the positive-sequence network. The sequence components of line-to-ground voltages at the generator terminals illustrate this configuration.
Zero-sequence current induces a voltage drop across the generator's neutral impedance and other...
Zero-sequence current induces a voltage drop across the generator's neutral impedance and other...
467
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
1.1K
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
1.1K


