Estimation of mutual information for real-valued data with error bars and controlled bias
Caroline M Holmes1, Ilya Nemenman2
1Department of Physics, Princeton University, Princeton, New Jersey 08544, USA.
Physical Review. E
|October 3, 2019
Summary
Estimating mutual information is challenging. This study enhances a popular method using neighbor distances, improving bias detection, variance estimation, and parameter selection for broader applications in complex systems analysis.
Area of Science:
- Information theory
- Complex systems analysis
- Statistical inference
Background:
- Estimating mutual information is crucial for analyzing complex, biological, and quantum systems.
- Existing estimators, including the Kraskov et al. method, face limitations in accuracy and applicability.
- Provably optimal estimators for mutual information do not exist, necessitating robust practical solutions.
Purpose of the Study:
- To address the limitations of existing mutual information estimators, particularly the Kraskov et al. method.
- To improve the accuracy, reliability, and applicability of mutual information estimation.
- To provide practical guidelines for using and validating mutual information estimates.
Main Methods:
- Critically evaluating the application of bootstrapping for variance estimation of mutual information.
- Developing an enhanced Kraskov et al. estimator with expanded applicability.
- Implementing methods for bias verification and variance estimation.
- Providing guidance on selecting the free parameter for the estimator.
Main Results:
- Demonstrated pitfalls of naive bootstrapping for mutual information variance estimation.
- Introduced an improved Kraskov et al. estimator with enhanced bias detection and variance estimation.
- Showcased improved performance on synthetic, neurophysiological, and systems biology datasets.
- Provided practical guidelines for parameter selection, enhancing estimator usability.
Conclusions:
- The enhanced mutual information estimator offers improved accuracy and reliability for complex systems.
- The developed methods provide a self-consistent framework for validating mutual information estimates.
- The study offers practical tools and guidelines for researchers in diverse scientific fields.
- The improved estimator is validated across various synthetic and real-world biological datasets.
Related Concept Videos
Random Error
7.6K
Random or indeterminate errors originate from various uncontrollable variables, such as variations in environmental conditions, instrument imperfections, or the inherent variability of the phenomena being measured. Usually, these errors cannot be predicted, estimated, or characterized because their direction and magnitude often vary in magnitude and direction even during consecutive measurements. As a result, they are difficult to eliminate. However, the aggregate effect of these errors can be...
7.6K
Statistical Analysis: Overview
14.0K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
14.0K
Estimating Population Mean with Unknown Standard Deviation
8.7K
In practice, we rarely know the population standard deviation. In the past, when the sample size was large, this did not present a problem to statisticians. They used the sample standard deviation s as an estimate for σ and proceeded as before to calculate a confidence interval with close enough results. However, statisticians ran into problems when the sample size was small. A small sample size caused inaccuracies in the confidence interval.
William S. Gosset (1876–1937) of the...
William S. Gosset (1876–1937) of the...
8.7K
Propagation of Uncertainty from Systematic Error
1.2K
The atomic mass of an element varies due to the relative ratio of its isotopes. A sample's relative proportion of oxygen isotopes influences its average atomic mass. For instance, if we were to measure the atomic mass of oxygen from a sample, the mass would be a weighted average of the isotopic masses of oxygen in that sample. Since a single sample is not likely to perfectly reflect the true atomic mass of oxygen for all the molecules of oxygen on Earth, the mass we obtain from this...
1.2K
Estimating Population Mean with Known Standard Deviation
9.5K
To construct a confidence interval for a single unknown population mean μ, where the population standard deviation is known, we need sample mean as an estimate for μ and we need the margin of error. Here, the margin of error (EBM) is called the error bound for a population mean (abbreviated EBM). The sample mean is the point estimate of the unknown population mean μ.
The confidence interval estimate will have the form as follows:
(point estimate - error bound, point estimate +...
The confidence interval estimate will have the form as follows:
(point estimate - error bound, point estimate +...
9.5K
Accuracy and Errors in Hypothesis Testing
545
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
545


