Related Experiment Video
Updated: Aug 7, 2025

Analysis of SEC-SAXS data via EFA deconvolution and Scatter
Published on: January 28, 2021
Exact Gaussian processes for massive datasets via non-stationary sparsity-discovering kernels.
Marcus M Noack1, Harinarayan Krishnan2, Mark D Risser3
1Applied Mathematics and Computational Research Division, Lawrence Berkeley National Laboratory, Berkeley, CA, 94720, USA. MarcusNoack@lbl.gov.
This study introduces a novel approach to Gaussian Processes (GPs) for large datasets. By enabling kernels to discover inherent sparsity, exact GPs can now scale beyond 5 million data points efficiently.
Area of Science:
- Computational Mathematics
- Machine Learning
- Scientific Computing
Background:
- Gaussian Processes (GPs) are powerful for stochastic function approximation, offering analytical tractability, robustness, and uncertainty quantification.
- Exact GPs face significant computational and storage challenges ([Formula: see text] and [Formula: see text] complexity) with large datasets.
- Existing approximate methods often compromise accuracy and kernel design flexibility.
Purpose of the Study:
- To develop a method for scaling exact Gaussian Processes to large datasets without compromising accuracy.
- To leverage naturally occurring sparsity in data and physical processes for efficient GP computation.
- To enable the design of flexible, expressive kernels that can learn sparse structures.
Main Methods:
- Proposed a novel approach where kernels discover, rather than induce, sparse structure.
- Developed ultra-flexible, compactly-supported, and non-stationary kernels capable of learning zero covariances.
- Combined kernel design with High-Performance Computing (HPC) and constrained optimization techniques.
Main Results:
- Achieved scalability for exact Gaussian Processes to datasets exceeding 5 million data points.
- Maintained analytical tractability and uncertainty quantification inherent to exact GPs.
- Enabled greater flexibility in kernel design for modeling complex physical processes.
Conclusions:
- This new method overcomes the computational limitations of exact GPs for large-scale applications.
- The ability of kernels to learn sparsity offers a more principled and accurate approach to GP approximation.
- This advancement significantly expands the applicability of exact Gaussian Processes in science and engineering.
Related Concept Videos
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Maxwell-Boltzmann Distribution: Problem Solving
This distribution function f(v) is defined by saying that the expected number N (v1,v2) of particles with speeds between v1 and v2 is given by
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
Extraction: Partition and Distribution Coefficients
For extracting a solute from an aqueous phase into an...
Distributions to Estimate Population Parameter

