Related Experiment Video
Updated: Nov 2, 2025

Stochastic Noise Application for the Assessment of Medial Vestibular Nucleus Neuron Sensitivity In Vitro
Published on: August 28, 2019
Doubly Stochastic Normalization of the Gaussian Kernel Is Robust to Heteroskedastic Noise
Boris Landa1, Ronald R Coifman1, Yuval Kluger1,2,3
1Program in Applied Mathematics, Yale University.
Doubly-stochastic normalization of affinity matrices robustly handles heteroskedastic noise in data analysis. This method, unlike others, accounts for varying noise levels, improving data point similarity assessments.
Area of Science:
- Data Science
- Machine Learning
- Computational Biology
Background:
- Affinity matrix construction is crucial for data analysis, often using Gaussian kernels on Euclidean data.
- Common normalization methods include row-stochastic and symmetric variants.
- Heteroskedastic noise, where data points have different noise variances, poses a challenge for these methods.
Purpose of the Study:
- To investigate the robustness of doubly-stochastic normalization for affinity matrices under heteroskedastic noise.
- To compare the performance of doubly-stochastic normalization against row-stochastic and symmetric normalizations in the presence of noise.
Main Methods:
- Constructing affinity matrices using Gaussian kernels on pairwise distances.
- Applying doubly-stochastic, row-stochastic, and symmetric normalizations.
- Theoretical analysis in a high-dimensional setting with heteroskedastic noise.
- Numerical simulations and analysis of single-cell RNA sequencing data.
Main Results:
- Doubly-stochastic normalization demonstrates robustness to heteroskedastic noise.
- Theoretical convergence rate of the noisy affinity matrix to its clean counterpart is m^(-1/2).
- Row-stochastic and symmetric normalizations perform unfavorably under heteroskedastic noise.
Conclusions:
- Doubly-stochastic normalization is advantageous for data with varying noise variances.
- This method offers improved exploratory analysis for datasets like single-cell RNA sequences.
- The findings highlight the importance of normalization choice in noisy, high-dimensional data analysis.
Related Concept Videos
Propagation of Uncertainty from Random Error
Chebyshev's Theorem to Interpret Standard Deviation
Propagation of Uncertainty from Systematic Error
Normal Distribution
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Random and Systematic Errors

